Cohere has released Parse 5, a 2.3-billion-parameter vision language model that converts PDFs, slides, and images into Markdown, extracting text, HTML tables, lists, forms, image descriptions, and bounding boxes in a single pass. Built on the North-Micro-Vision-Instruct architecture, Parse 5 supports nine stable languages, offers an 8,192-token context window, and is generally available via API, Microsoft Foundry, AWS SageMaker, and single-tenant Model Vault deployments. Cohere prices the Parse API at $1.50 per 1,000 pages, with higher-volume Model Vault instances for dedicated capacity, and reports a 79.2 ParseBench subset score that customers are encouraged to validate on their own documents.
This update represents a notable development in the Ai sector. Organizations and founders tracking this space should evaluate potential strategic and technical implications on their operations.