OCR for documents already sitting in your own storage.
Point it at an S3 prefix, get Markdown and JSON back in the same bucket
under _ocr/. Your files never leave your account.
Upload one page and see for yourself, or read the pricing.
Markdown, real <table> HTML with rowspan and colspan intact,
and every region typed with its pixel box. Measured at 0.912 TEDS
on OmniDocBench — level with the published state of the art.
Name the fields you want and get them back as key/value pairs:
1 as lastname, [due date] as due, or
if it is a driver license: …. Included on every plan.
Read from your prefix, write results beside the source. We hold nothing after the job finishes — credentials are dropped the moment it ends.
A queue with per-key fairness, so one large job cannot starve a small one behind it.
PDFs are rasterised a page at a time, so a 500-page file does not need 500 pages of memory.
_ocr/<key>.md and
_ocr/<key>.json back to your bucket.Access is currently invite-only while capacity is limited.