Skip to content

Document transcription · Data extraction

We turn documents into accessible information.

Handwritten, printed and typewritten, from any period and in any condition, turned into information that can be searched, filtered and cited.

f. 1v · The problem

Digitized doesn't mean usable.

Millions of documents have been digitized, and we're still working with them as if they were paper. An image of a page preserves the original, but it doesn't let you search for a name, filter by date, count how many times a term appears, or cross-reference one file with another.

The cost doesn't show up in the budget, it shows up in time: inquiries that take weeks to answer, research limited to what someone managed to read, and entire files nobody opens because opening them takes years.

It isn't a preservation problem. It's an access problem.

f. 2r · Document transcription

We read and transcribe your documents.

From a scan or photograph of each page, we get the full text. We work with material from any period and in any condition, including deteriorated manuscripts and printed material, in any language. Our AI models adjust to the type of handwriting and the period of each document.

From the image to text you can search.

f. 2v · Data extraction

We extract the data you need from each document.

The models identify dates, names, amounts and file numbers within the text. We define them together with you based on what you need to do, and you receive them in a table you can sort, filter and cross-reference with other sources.

Fields we extract

dates · names of people · places · occupations · amounts · file or folio numbers · document type · signatories · cross references

We structure all your data into one table.

f. 3r · AI Agent

An AI agent that knows all your documentation.

We train an artificial intelligence agent on the documents already processed. It answers specific questions in natural language and cites the document and page each answer comes from, so everything can always be checked against the original.

AI agent for your documentation 1,240 pages

Cited sources

Get answers in seconds.

f. 3v · Delivery

Delivery and integration of the data

Besides delivering files, we bring the information into the system your organization already works with day to day.

Files

Delivery formats

  • Plain or structured text, by page and by document
  • Searchable-text-layer PDF, preserving the original image
  • Spreadsheet with the extracted fields
  • Standard document-description formats
  • Coverage report on the processed batch
Systems

Integration with existing systems

  • Direct load into a database or document management system
  • Publication in a repository or consultation portal
  • API connection to software you already have installed
  • Periodic batch delivery, when material arrives on an ongoing basis

f. 4r · Modes

What's your case?

There are two ways of working with us, with different timelines, scope and terms.

A specific file or batch

A bounded set: a file, a series, a box or a group of records. It gets processed in full and you receive the data. It has a beginning and an end.

  • Quote calculated from a sample of the material
  • Delivery within days, or in stages when volume requires it
  • One-time payment
Send a sample

Continuous incoming documentation

The same type of document received periodically, or a large file that must be processed in full and kept up to date. We build the process into your organization's operations.

  • Measured trial before scaling up
  • Integration with existing systems
  • Scope and budget defined from the trial
Request a meeting

f. 4v · Process

The process, step by step

01

Sample

You send us a sample of your material. Between twenty and fifty pages are enough to assess the material and its difficulty.

02

Material preparation

Rotation, contrast, noise and cropping are corrected. A significant part of the final accuracy is determined at this stage, before the content is processed.

03

Processing

The material is transcribed with artificial intelligence models tuned to the document type, the handwriting and the period, and the previously defined fields are extracted, each one linked to its location on the original.

04

Delivery and integration

You receive the files in the agreed format and, where applicable, the data loaded directly into your system, along with the batch's coverage report.

f. 5r · Sectors

Who works with us

Archives, libraries and museums
to open their files up to consultation and research.
Universities and institutes
to transcribe complete document series and make them citable.
Public administration
to turn digitized files into searchable information.
Law and notarial practices
to locate information within deeds, protocols and case files.

When document work needs to become a system

Beyond transcription and extraction, we build AI agents trained on your organization's information, integrations with existing systems, and custom software.

See all services

f. 5v · Questions

Frequently asked questions

Does it work with old handwriting?

Yes, it's the case we work with most. Old-style handwriting, period abbreviations and obsolete spelling are adjusted to the period of each document.

What languages do you work in?

Spanish, English, Catalan, French, Portuguese and Italian, including their older forms. For material in other languages, it's best to check with us using a sample.

How long does a large file take?

Processing itself is fast; the timeline is mostly determined by material preparation. A batch of a thousand pages is usually delivered within days. A file of tens of thousands is planned in stages, so the first results can be used before the project is finished.

What happens with our documents?

Processing can run on our infrastructure, on yours, or by installing the process on your own servers, so the material never leaves your organization. Before starting, we agree in writing on where it's processed, who has access, and what happens to it once it's finished. If your institution has specific data-handling requirements, it's best to raise them in the first conversation.

What if the scans are in poor condition?

Pre-processing for rotation, contrast and noise significantly improves the result. If the material still doesn't allow for usable results, we let you know at the sample stage, before committing to the full project.

We haven't digitized yet. Can you advise us?

Yes. We can guide you on how to digitize so that later processing works correctly: resolution, lighting, page order and file naming. These decisions seem minor but shape the entire project.

Do you take part in funded projects?

Yes. We can join as a technical partner in applications for research, digitization or digital humanities funding, and provide whatever technical documentation the project requires. It's best to contact us with the call's deadlines in mind.

What AI technology do you use?

Handwriting-recognition models, both manuscript and print, combined with page-layout analysis: margins, columns, tables and marginal notes. That lets a ruled register or a form with fields be processed while respecting its structure, instead of turning into one continuous block of text. The models also adjust to the particularities of each collection: the dominant handwriting, its own vocabulary and period abbreviations. That adaptation is the difference between an acceptable result and a usable one.

What level of accuracy can I expect?

It depends on the material, which is why the work always starts with a sample: on it, we deliver the accuracy rate obtained with your own documents and the characteristics that affect that result. With that information you can decide whether to move forward before committing to a budget.

f. 6r · Contact

Send us a sample and we'll show you what can be obtained from it.

Before any commitment, we process part of your material and deliver the data along with the accuracy obtained on your own documents.