How Do XML Standards Help Teams Manage Complex Documentation?

How Do XML Standards Help Teams Manage Complex Documentation?

By their very nature, very large documentation sets covering thousands of topics in hundreds of products, written in scores of languages for scores of countries and regulatory bodies, are prone to certain types of failure. The principal failure in large, complex, information-rich systems is that content is copied rather than reused. Formatting is also locked into the source rather than left to the rendering system. There is no guarantee of what any given safety warning actually looks like when it is published.

Separating what content means from how it looks

At its core, the rendering of semantic source XML files is simply based on the designated semantic roles of the contained content (prerequisites, steps, warnings, parameter names, etc.) as opposed to visual appearance. The appearance itself is then left to stylesheets and the processing of the said source files in due course.

This is particularly valuable for distinguishing a writer’s content from the presentation. Instead of having a writer accidentally change the appearance of a caution statement in a single manual, the appearance of that statement is defined elsewhere. Thus, a single change in a brand or in a regulatory presentation requirement can be made in one place and that change will appear in all documents that use that definition.

Validation as a quality gate

Schemas and DTDs can also enforce editorial rules in a very strict way. For example, if all your procedures are supposed to start with a purpose, the schema can be set up to refuse documents that are missing a purpose statement. This kind of enforcement is more effective the more you have written by engineers and subject specialists rather than dedicated technical writers.

Modular authoring and the economics of reuse

A topic is a self-contained unit of work. A 400-page manual consists of hundreds of topics which are organized by a map or manifest to form deliverables. Thus, an installation topic can be used in a quick start guide, a full service manual and a knowledge base article without having to create duplicates of it.

Reuse at these different levels is also possible but typically only the first level of reuse is implemented before the team runs out of steam.

  • Topic reuse, where an entire module is referenced by multiple publications.
  • Fragment reuse, where a warning, a specification table, or a legal clause is pulled into many topics by reference.
  • Variable substitution, where product names, version numbers, and part codes resolve at build time.
  • Conditional filtering, where a single source produces audience-specific or market-specific outputs from profiling attributes.

This reuse also compounds very positively for translation, because a given piece of content is translated only once, not once for each different output in which it appears. Only changes to a previously translated piece of content will require re-translation for later releases. For a company publishing in twelve languages, this is usually where the greatest business value is found.

Choosing a standard that matches your content

Document standards are not equal; each is defined to cover a specific set of information, and targets a specific audience. You are actually creating years of headaches by choosing one simply because it’s familiar.

 

Standard Typical use Main strength Main constraint
DITA Software, hardware, and product documentation Mature reuse model, wide tool support, specialization Learning curve for maps and keys
S1000D Aerospace, defense, and complex assets Data module coding tied to the physical breakdown Heavy process overhead for smaller programs
DocBook Technical books and long-form references Rich narrative structures, stable for decades Weaker topic-level reuse
JATS Scholarly and research publishing Metadata and citation fidelity Narrow applicability outside publishing

Specialization rather than invention

Most teams will need to customize the base set of topics a standard provides, often constraining or even specializing the standard to be able to transform the resulting documents to standardized output, and to use tools already in use by other teams. If your content is product-centric, the DITA XML standard is usually the safest foundation to specialize from, since it already anticipates that kind of extension. Writing a custom schema from scratch is fast and powerful for short periods of time, but creates problems down the road that are harder to solve than if a standard had been used in the first place.

Publishing from one source to many outputs

Once you have valid, modular content, generating different types of output such as PDF for government filing, HTML for web sites, and structured data for products such as in-product help or chatbots becomes a straightforward build process.

Run all of the normal validation (e.g. Links) as part of your normal continuous integration. Tag off released versions of content to specific versions of your repository as you would with code. Use that to serve up off-line content, to answer historic questions in your audits.

Making the transition without stalling delivery

Conversion of legacy documentation to structured content fails when the entire back catalog is migrated through a single conversion project. Instead, keep producing output while developing the conversion tool chain.

  1. Audit existing content for genuine duplication and identify the highest-value reuse candidates.
  2. Define the information model and authoring rules before buying tooling, not after.
  3. Convert one active product line, publish it through the new pipeline, and fix what breaks.
  4. Migrate legacy material only where it is still maintained, and retire the rest.
  1. A) Plan the costs of authors, reviewers and a model owner for the necessary work as thoroughly as you would for the needed software. The goal of authors is to write standalone topics. The goal of reviewers is to approve fragments of content rather than completed pages of content. The living information model needs to have a model owner.