Unlocking archives through intelligent automation.
From image to insight — Osiris transforms complex historical documents into structured, searchable data.
Osiris develops advanced tools that transform complex archival and historical materials into structured, searchable data. Our technology automates processes that once required extensive manual work — from page segmentation and handwriting recognition to data extraction from complex layouts — enabling faster, more accurate research and digital preservation.
Who we are.
Oliver Buxton Dunn
Head of Research and DevelopmentAlexis Litvine
Head of company operationsYiannos Stathopoulos
Co-founder and tech advisorRuth Murphy
Ruth holds a PhD in Italian from the University of Cambridge and is currently a postdoctoral researcher at the University of Sheffield. As a talented linguist, she oversees modern transcription work for Osiris.
Chloe Ashley
Chloe Ashley helps with company administration and client care. She also lends a hand with the more delicate transcription tasks and data QA. She holds a degree in English Literature and Hispanic Studies from Queen Mary, University of London.
Stan Hinton
Stan graduated in history from Cambridge after a successful career start in the tech industry. He specialises in UI components and annotations schemes for AI.
Charlie Cook
Charlie is an undergraduate mathematics student at The University of Aberdeen with an enthusiasm for coding and AI. Charlie is involved in model training and fine tuning, as well as script construction and refinement.
Sam Hoyle
Sam is now ESRC-funded postgraduate student at the University of Durham, specialising in maritime violence in the Indian Ocean. He has helped us with transcription, segmentation, and image processing since 2023.
Why specialist document AI matters for historical collections
Documents can now be uploaded to frontier AI models such as Gemini, Claude, GPT, Qwen, Llama, Mistral, and other multimodal systems, and they can return surprisingly detailed results. These models are extremely powerful. We use them ourselves.
But historical collections are different from ordinary modern documents.
They are often difficult because of their language, layout, handwriting, abbreviations, damaged paper, faded ink, marginal notes, complex tables, obsolete terminology, multilingual content, and inconsistent structure. A model may produce an impressive-looking result while missing important details, misunderstanding the source, applying the wrong structure or "corrections", or failing silently on the parts of the document that matter most.
Successful document AI is therefore not just a question of choosing a powerful model. It depends on knowing the source material, selecting the right technology, preparing the images correctly, combining OCR, HTR, layout analysis, visual language models, validation tools, and structured extraction pipelines, and designing the workflow around customer data requirements - precisely.
This is what Osiris does.
We combine historical source knowledge with machine learning development and research-oriented production. We help clients choose the right models, adapt them where needed, and build the data pipeline required to produce reliable outputs at scale.
The AI model landscape is also changing rapidly. Hugging Face now hosts millions of models, with specialist communities forming around particular domains, languages, and problem types. The challenge is no longer simply finding a model; it is knowing which model, workflow, and validation process is right for a specific collection. Osiris helps clients navigate this landscape, select appropriate models, fine-tune them where required, and build a model and workflow that just works.
What Osiris provides
Historical source knowledgeWe work with complex archival and historical material, including handwritten records, early printed texts, tabular sources, administrative documents, registers, newspapers, maps, forms, and multilingual collections.
Model selection and developmentWe assess which models and tools are appropriate for each collection, including OCR, HTR, layout analysis, multimodal AI, and specialist open-source or commercial models. Where required, we train or fine-tune models for the script, layout, language, and structure of the source.
The model you needOff-the-shelf tools are useful, but many collections require a bespoke or adapted workflow. We build the model and pipeline needed for your documents, your accuracy requirements, and your final data use.
End-to-end productionWe manage the full process from image preparation and layout analysis through to OCR/HTR, structured data extraction, validation, reporting, and export.
Data pipelines, not one-off promptsWe do not simply upload documents to a model and accept the answer. We combine models with preprocessing, segmentation, batching, extraction logic, validation, review tools, and export formats designed for production use.
Trustworthy reportingWe identify where the workflow performs well, where errors are likely, and what level of review or correction is needed for the intended use. This helps isolate strengths and weaknesses before and during large-scale processing.
Privacy and data sovereigntyWhere required, we can process collections from start to finish on local or private servers, avoiding unnecessary data transfer to third-party AI platforms and reducing the risk of data leakage.
Outputs designed for useWe deliver data in formats suited to search, research, publication, discovery systems, databases, maps, catalogues, or customer ingestion pipelines.
Our multi-stage pipeline for document understanding.
Osiris is not a single LLM system. It is a multi-stage pipeline combining image preparation, custom HTR/OCR, segmentation, structured extraction, validation, and feedback loops to process complex documents reliably at scale.
Each stage is engineered to address specific failure modes — preserving structure, reducing hallucinations, and verifying outputs against the source. This allows us to deliver accurate, auditable results on material where generic AI tools break down.
Image prep
Sampling, de-warp, de-skew, crop, and enhancement for cleaner inputs.
HTR / OCR fine-tuning
Assisted training around your collection and document patterns.
Pre-segmentation
Split pages into useful zones, records, and fields before extraction.
Structured extraction
Targeted prompts, smaller inputs, deterministic checks, and review-ready outputs.
Structured output
Corrected text, structured data, and error metrics. We can also provide readable formats like PDF.
Accuracy you can measure — and trace.
General-purpose LLMs can transcribe reasonably well — but they also invent plausible, correct-looking text. Osiris measures and controls this risk explicitly.
Osiris gets the best from fast-evolving AI through research and development to provide trustworthy and useful information from complex records.
We regularly benchmark our models against competing systems (including industry leaders). The examples below report overall accuracy (F1) against ground truth and break errors down by type — isolating hallucination risk, numeric failures, and OCR noise.
18th-century English handwriting
HTR benchmark against ground truth.
Overall Accuracy (F1) — 18th C English handwriting Osiris ████████████████████ 90% ChatGPT █████████████████ 79% Error Types (Extra Tokens vs GT) Plausible errors (easy to miss) Osiris ██████ 3.97% ChatGPT ██████████████████████ 14.67% Numeric errors (high impact) Osiris █████████████████ 4.30% ChatGPT ██████████████████████ 5.67% Minor OCR errors Osiris ██████████████████████ 1.66% ChatGPT ██████████████████ 1.33% Other major errors Osiris ██████████████████████ 0.66% ChatGPT 0.00%
English language newspaper print (1925)
OCR benchmark against ground truth.
Overall Accuracy (F1) Osiris ████████████████████ 92% ABBYY ████████████████ 71% ChatGPT ███████████████████ 85% Error Types (Extra Tokens vs GT) Plausible errors (easy to miss) Osiris ███████████████ 4.66% ABBYY █████████████████ 5.27% ChatGPT ██████████████████████ 6.78% Numeric errors (high impact) Osiris ███████ 0.62% ABBYY ██████████████████████ 1.86% ChatGPT ███████ 0.58% Minor OCR errors Osiris █ 0.52% ABBYY ██████████████████████ 14.68% ChatGPT ██ 1.05% Other major errors Osiris ████ 1.35% ABBYY ██████████████████████ 6.83% ChatGPT ██ 0.58%
Overall accuracy is reported as F1 against ground truth. Error categories isolate hallucination risk, numeric failures, and OCR noise — helping teams prioritise what to check and what to fix.
Past projects.
Trusted by
Get in touch with us ...
Whether you need to extract data at scale, want to integrate geospatial data and historical records or need data consultancy, we'd love to hear from you.



