Legacy Content · XML & JSON Conversion · AI-Enabled · Accessibility & Metadata

From legacy formats to structured, reusable content.

We convert books, journals, educational materials, and legacy archives into the format your workflow requires — JATS, BITS, NLM DTD, DocBook, DITA, QTI, JSON, and more — using AI-assisted processing, with accessibility tagging and metadata enrichment built in.

25+ years in publishing technology
1M+ pages converted a year
Trusted by publishers worldwide Wolters Kluwer Pearson Informa Wiley Britannica McGraw-Hill Ingenta Edanz ASCO RCNI Mondadori

What each format enables

Choose the format your workflow needs

Each format we convert to opens up a specific set of platforms, workflows, and opportunities. Here's what becomes possible once your content is properly structured.

For journal publishers
  • JATSContent indexed by PubMed, Scopus, and Web of Science, and delivered to platforms like HighWire, Silverchair, and Atypon.
  • NLM DTDLegacy articles integrated with older publishing systems and submitted to PMC without requiring a full platform migration.
For book publishers
  • BITSBooks and monographs made discoverable on academic platforms, library systems, and digital repositories.
  • DocBookA single source file that publishes to print, PDF, HTML, and EPUB — without reformatting each output separately.
For education & assessment
  • DITAContent modularised so individual topics can be reused across courses, guides, and products — reducing the cost of updates significantly.
  • QTIAssessments and tests made portable across any LMS, assessment platform, or content repository without rebuilding them each time.
For digital platforms & AI
  • JSONContent consumable by web platforms, APIs, recommendation engines, and AI applications — the format modern systems expect natively.
  • CustomDirect integration with your platform or aggregator’s proprietary schema, with no additional transformation layer required.

Proof, not promises

A recent JATS conversion at scale

40K+Articles converted

Legacy articles brought up to JATS — and made accessible

A global academic content platform serving 1.6M+ readers across 195 countries, working with 130+ publishers, needed older articles brought up to JATS so they’d meet indexing requirements. We mapped a transformation roadmap, built custom XSLT transforms, converted 40,000+ articles to JATS-compliant XML, and converted select content to NIMAS for accessibility — then republished it directly into their production pipeline.

Read the full case study →

How it works

Five steps from first call to live output

No long procurement process before you see a result — the sample batch comes before any commitment.

01

Discovery call

Tell us your source format, schema, and indexing requirements.

02

Schema agreement

We confirm the target schema — JATS, BITS, NLM DTD, DocBook, DITA, QTI, JSON, or any custom format — and any publisher-specific customisations.

03

Free sample batch

5 articles or 2 book titles, converted at no cost, in 5 working days.

04

Full project

Your backlist converted at an agreed turnaround and price per page.

05

Delivery

Validated output delivered into your platform or repository.

How AI fits into the workflow

Smarter processing, human where it matters

We use AI-assisted analysis to identify document structure, tables, equations, references, and metadata — then route content through the right processing path based on complexity. Validation and enrichment are automated where repeatable; human reviewers handle exceptions and edge cases.

  • AI-assisted element identification — structure, metadata, equations, and references flagged automatically
  • Complexity-based routing — simple, medium, and complex content processed through the right workflow
  • XML & JSON mapping — to JATS, BITS, NLM DTD, DocBook, DITA, QTI, JSON, and custom schemas
  • Automated validation + expert QA — machine checks first, human judgment where it matters most

Why publishers choose our approach

Built for reuse, not just delivery

AI where it adds value

Automation handles repeatable work. Experts focus on exceptions, edge cases, and output quality.

Flexible standards

JATS, BITS, NLM DTD, DocBook, DITA, QTI, JSON, and custom schemas — we work to the format your workflow requires.

Accessibility-aware

Structured content preserves and enriches the information needed for accessible digital outputs as standard.

What we handle

Built for the content that's hard to convert well

Schemas & standards

We work to the schema your indexing service or repository requires — JATS, BITS, NLM DTD, DocBook, and publisher-specific DTD customisations.

JATS BITS NLM DTD DocBook

Source formats

Word, InDesign, born-digital and scanned PDF, LaTeX, and legacy XML in custom schemas all come into the same pipeline.

.docx InDesign PDF / OCR LaTeX

Content others find difficult

Equations in MathML, chemical structures, multilingual text, complex tables, figures with captions, footnotes, and cross-references are handled as standard, not as exceptions.

Validation, on every file

Schema and Schematron rule validation, custom XSLT checks, and a manual editorial QA pass — before anything is delivered.

Tooling

oXygen XML Editor, XSLT 2.0/3.0 transforms, Python-based automation, and custom CMS integrations for repeat batches.

Speed & capacity

Our team converts up to 35,000 pages a month, across journals, books, educational content, and legacy archives.

See your own content converted — free

Send us 5 articles or 2 book titles. We'll convert them to your schema in 5 working days, at no cost, with no commitment.