Skip to content

Recognise text ​

Turns an image of handwritten or printed text into a transcription with coordinates for every region, line and word.

What it does ​

Recognizes the text on a single document image using a Transkribus model, returning the transcription together with the layout: regions, lines, baselines, coordinates and, where available, words.

Choosing a model ​

Two families of model do this, and which you pick matters more than any parameter. See Model types for the full picture:

  • Super models are large general-purpose models that work across many hands, languages and centuries without any training. Use these when you have mixed or unfamiliar material, or no ground truth. They are the right default.
  • PyLaia models are small, fast and trained on a specific hand, collection or script: considerably cheaper at volume, and a well-trained one beats a general model on its own material.

Both are addressed the same way, by model ID (htrId). See Models for how to find one.

Input ​

A single image (JPEG, PNG or TIFF, up to 20 MB), as a base64-encoded string or an http(s) URL, plus the htrId of the model to apply. Optional: pre-existing layout information and layout model parameters via config.lineDetection (the field name is historical; these are layout models producing baselines, regions and reading order).

Output ​

The recognized text and layout, available inline in the job's content field, as a ZIP archive, or as PAGE XML (PRImA 2013-07-15 schema). See the Job schema in the API Reference.

Parameters ​

config.textRecognition.htrId (required), config.textRecognition.languageModel, config.lineDetection.*. Full schema and constraints are in the API Reference (JobRequest, JobConfig, TextRecognitionConfig, LineDetectionConfig).

Cost ​

Credits at 50% of the rate of the same job in the Transkribus platform. See Credits and cost.

Limits ​

See Rate limits: 20 MB per image, one image per job, results retained 24 hours.

Transkribus Developer Platform