# Mindee Documentation

Mindee helps anyone automate document processing through data recognition, computer vision, and machine learning.

Our mission is to provide real-time, human-level accuracy in data extraction from paper and digital documents.

<a href="https://app.mindee.com/signup?utm_source=docs" class="button primary" data-icon="user-plus">Sign Up to Mindee</a>

## Discover Mindee

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-cover data-type="image">Cover image</th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><mark style="color:red;"><strong>Quickstart</strong></mark></td><td>Get started in minutes with our step-by-step guide.</td><td><a href="/files/3kszqAd8yughVKHMV6mI">/files/3kszqAd8yughVKHMV6mI</a></td><td><a href="/pages/A5YeQZj3V1lPs3aUBvyZ">/pages/A5YeQZj3V1lPs3aUBvyZ</a></td></tr><tr><td><mark style="color:red;"><strong>Features</strong></mark></td><td>Explore key capabilities and how they solve your document challenges.</td><td><a href="/files/YXF5QP4ZHwLkmWrqc5y8">/files/YXF5QP4ZHwLkmWrqc5y8</a></td><td><a href="/pages/8OxOdj9XqhGhCUshcVID">/pages/8OxOdj9XqhGhCUshcVID</a></td></tr><tr><td><mark style="color:red;"><strong>Live test</strong></mark></td><td>Upload a sample document and see the extraction in action.</td><td><a href="/files/EqAaU2oEWaIPuv9irLEe">/files/EqAaU2oEWaIPuv9irLEe</a></td><td><a href="/pages/biWdfc0r8eaL9bD8oVyh">/pages/biWdfc0r8eaL9bD8oVyh</a></td></tr><tr><td><mark style="color:red;"><strong>Integration</strong></mark></td><td>See how you can easily connect Mindee with your stack.</td><td><a href="/files/bcnF8PcHcoCxh6q7ZweF">/files/bcnF8PcHcoCxh6q7ZweF</a></td><td><a href="/pages/iREYUuRlIfMqqSxTfHZF">/pages/iREYUuRlIfMqqSxTfHZF</a></td></tr><tr><td><mark style="color:red;"><strong>SDKs</strong></mark></td><td>Official Client Libraries (SDKs) for Python, Node.js, and other languages to plug Mindee into your stack.</td><td><a href="/files/12COQpY2wUXxA881jlOW">/files/12COQpY2wUXxA881jlOW</a></td><td><a href="/pages/vw7TptFVGhNOzfQ8n5MH">/pages/vw7TptFVGhNOzfQ8n5MH</a></td></tr><tr><td><mark style="color:red;"><strong>Plans</strong></mark></td><td>Review pricing and choose the plan that fits your usage.</td><td><a href="/files/Mzty1XNWzr5Ayqu0MAsc">/files/Mzty1XNWzr5Ayqu0MAsc</a></td><td><a href="/pages/DPTbMyyH4dy23l2YrYz7">/pages/DPTbMyyH4dy23l2YrYz7</a></td></tr></tbody></table>

{% hint style="info" %}
**This is the documentation for:** [**app.mindee.com**](https://app.mindee.com/) **(V2 - Latest)**

If you're looking for the documentation for: [platform.mindee.com](https://platform.mindee.com/) (V1)

Take a look at the [V1 Section](/v1)
{% endhint %}


# Key Features

Mindee's key features.

Mindee headline features center on custom models, RAG, confidence scores, and rich outputs.

## Custom Extraction Models

Build fully customizable extraction models with interactive data schemas and an AI assistant for fast iteration and control over fields and guidelines.\
Learn more at [Data Schema Overview](/extraction-models/data-schema).

Live Test UI for side-by-side visual validation, JSON inspection, and RAG simulation before production.\
Learn more at: [Live Test](/models/live-test).

## Extraction Model Catalog

Choose from our model templates with predefined fields for data extraction.\
You can modify models created from templates as needed, just like custom models.

Some popular use cases include:

* **Billing**: Financial Document, Invoice, Receipt, ...
* **Identification**: Passport, ID Card, Driver License, ...
* **Employment**: Resume or CV, Payslip, ...
* **Banking**: Check, Bank Statement, ...
* **Logistics**: Bill of Lading, Vehicle Registration, ...

Learn more at: [Extraction Use Cases](/use-cases/extraction-models)

## Utility Models

In addition to extraction models, Mindee also provides utility models that can be used either separately or in conjunction with Extraction Models.

These utilities allow for more complex workflows and greater flexibility.

* **Split**: Detect separate documents in a multi-page source file, matching each document to a category.
* **Crop**: Detect document boundaries on each page, matching each document to a category.
* **Classification**: Automatically assign images and scanned documents to the right category.
* **OCR**: Extract raw text and word boundaries from any document with high precision.

{% hint style="info" %}
Utilities are currently in beta: they are usable, with full integration support coming soon.
{% endhint %}

## Advanced Features

Retrieval-Augmented Generation (RAG) to continuously improve accuracy using your own documents and corrections.\
Learn more at: [Continuous Learning (RAG)](/extraction-models/optional-features/improving-accuracy).

Confidence scores per field to boost precision and drive downstream decisions. Allows automating your workflows.\
Learn more at: [Confidence Score and Accuracy Boost](/extraction-models/optional-features/automation-confidence-score).

Polygon coordinates for every extracted field to map values to document regions.\
Learn more at: [Polygons (Bounding Boxes)](/extraction-models/optional-features/polygons-bounding-boxes).

Full text content output for complete OCR alongside structured fields.\
Learn more at: [Raw Text (Full OCR)](/extraction-models/optional-features/raw-text-full-ocr).

Data residency controls (EU/US) for privacy and compliance (GDPR, SOC 2 Type II).\
Learn more at: [Data Processing Policies](/models/data-processing-policies).

## Integrations

Asynchronous processing with polling and webhooks; official SDKs for Python, Node.js, Ruby, PHP, Java, and .NET.\
Learn more at: [Client Libraries / SDKs](/integrations/client-libraries-sdk).

No‑code and low‑code integrations, including an official Make.com app and generic HTTP workflows.\
Learn more at: [No-Code Integration](/extraction-models/no-code-integration).

## Flexible Plans

Mindee pricing is plan-based with included pages, overages, and feature tiers.

Pay annually to save money, or monthly.

Compare all plans and their features at: [Plans and Credits](/account-management/plans).

<a href="https://app.mindee.com/signup?utm_source=docs" class="button primary" data-icon="user-plus">Sign Up to Mindee</a>

## Need Something Else?

Our product team is always interested in adding new features to help our users.\
[Make a feature request!](https://feedback.mindee.com/?b=682f69c9e2404756e7e68d1c)


# Quick Start

Follow the fastest path from account creation to your first successful document processing request with Mindee.

## Sign Up!

<a href="https://app.mindee.com/signup?utm_source=docs" class="button primary" data-icon="user-plus">Sign Up to Mindee</a>

You can sign up and connect using your Google or Github account.

You can also create an account using good old-fashioned email and password.

## Create a Model

The quickest way is to create a model from a template in the catalog.

If there are no model templates that exactly fit your needs, you can:

* [Modify a model from the catalog](/getting-started/defining-a-model#from-an-existing-model)
* [Create your own model from scratch](/getting-started/defining-a-model#from-scratch)

## Test the Model

Be sure to make at least one test on a document, using the [Live Test feature](/models/live-test).

If it doesn't work as expected, modify the Data Schema, as explained in: [Models Overview](/models/models-overview#modifying-your-data-schema).

## Create an API Key

On the left-hand menu of the Mindee Platform, click [**API Keys**](https://app.mindee.com/settings?tab=api-keys) and create a new key.

## Send a File or URL

You can easily send a file or an URL using one of our [officially-supported SDKs](/extraction-models/sdk-integration).

For a quick introduction and ready to use code samples, check: [Extraction Quick Start](/extraction-models/sdk-integration/quick-start).

## Process the Result

You'll need some way of importing the JSON structure and manipulating it.

For quick testing, we recommend to using the [CLI feature](/integrations/client-libraries-sdk/command-line-tools-cli) of our client libraries.

{% @supademo/embed url="<https://app.supademo.com/demo/cmewtri4u8939v9kqh3q6u1g7>" demoId="cmewtri4u8939v9kqh3q6u1g7" %}


# Defining a Model

Build your model from a prebuilt template or create one from scratch using a sample document.

## From a Model Template

Chances are, there is already a model template that at least partly fits your needs.

Click on Catalog to check the model templates, and when clicking on one of them, you'll create a new model based on the template.

You now have your own model, you can start testing it immediately.

If you need to make adjustments, use the "Data Schema" tab as explained in more detail in: [Models Overview](/models/models-overview#modifying-your-data-schema).

For instance, here is a demo of how to create a model from the Invoice model template:

{% @supademo/embed url="<https://app.supademo.com/demo/cmieehzyjar67b7b4lf5tgvrb>" demoId="cmieehzyjar67b7b4lf5tgvrb" %}

## From Scratch

If you don't find something that matches your needs in the catalog, you also have the possibility to click on **Create Custom Model.**\
\
You'll have two possibilities :

* upload a sample document
* write the document type

Our AI Agent will then create a custom model for you.\
\
You now have a model with a **custom** schema that you can modify, using the "Data Schema" tab as explained in more detail in: [Models Overview](/models/models-overview#modifying-your-data-schema).\
\
Here is a tutorial that shows how to create a custom model by filling in the document type and by uploading a sample doc.

{% @supademo/embed url="<https://app.supademo.com/demo/cmiem2hsob2x5b7b4ye1il6a8>" demoId="cmiem2hsob2x5b7b4ye1il6a8" %}


# Glossary

A glossary of common Mindee terminology.

You will frequently encounter these concepts throughout this documentation as Mindee is centered around them.

### Annotation

The values (labels) a human enters to train, test, or alter an [ML](#ml-machine-learning) processing pipeline.

Mindee uses annotations notably for the RAG feature.

### API (Application Programming Interface)

A set of protocols and rules to define how applications or devices can connect to and communicate with each other.

In the context of Mindee, the API defines how to send files and receive results.

### Catalog

A list of [model](#model) templates you can search and use to create your own models.

### Custom Model

A [model](#model) you create from scratch for your own [document](#document) type and use case.

### Data Schema

A set of [fields](#field) and their [guidelines](#guideline) for data extraction. Associated to a [model](#model).

### Document

A structured or semi-structured written or graphical representation of human thought, such as an invoice, receipt, ID card, W9 form, contract, train ticket, etc.

Mindee processes documents contained in various [file](#file) formats.

### Extraction Model

A type of Mindee model designed to extract structured data from a [document](#document).

Includes a [Data Schema](#data-schema) and any optional features.

### Field

A single data point to extract from a document. Part of a [Data Schema](#data-schema).

### File

A computer file. In the context of Mindee, a file may contain one or several [documents](#document).

### Guideline

A set of instructions provided by the user, to clarify or alter the behavior of a [ML model](#ml-machine-learning).

### Inference

The processing step where a [ML model](#ml-machine-learning) performs an operation on a [file](#file).

By extension, this also refers to response or results that our API returns for a given file.

### Job

The asynchronous processing task created when a file or URL is sent to the Mindee [API](#api-application-programming-interface) for processing.

### ML (Machine Learning)

Algorithms that learn and generalize from data, and can perform tasks without explicit programming instructions.

This is the backbone of Mindee's technology.

### Model

A set of instructions for processing [files](#file). There are several types of Mindee models that correspond to different uses.

A model can operate on a single document or multiple documents within the same file.

### Multimodal Model

A type of deep-learning model that processes multiple types of data, such as text and vision.

Mindee uses multimodal models when processing documents.

### OCR

Optical Character Recognition (OCR) is the process by which a machine reads text on printed documents and transcribes them into a machine-readable format, like a `.txt` file.

### OpenAPI

An industry specification that defines a language-agnostic interface for RESTful APIs.

### Payload

A payload refers to data that is submitted to the Mindee server and the data returned by the server when an API request is made.

### Polling

A result retrieval flow where your application repeatedly checks the processing status of a [job](#job).

### Mindee

Pronounced "mind - ee", because we are like a helpful **mind** to you!

Mindee is the leading developer tool that powers document understanding for top players and unicorns like Qonto, Spendesk , or Payfit.

Founded in 2018, we raised a $14M Series A led by GGV Capital in 2021, just after graduating from Y Combinator (YC).

### SDK (Software Development Kit)

A set of software development tools developers can use to facilitate the creation of applications.

### Utility Model

A [model](#model) used to preprocess or analyze documents, such as Split, Crop, Classification, or raw OCR.

### Webhook

A result retrieval flow where Mindee sends the processing result directly to your Web server.


# Models Overview

Understand the core concepts of Mindee models, the different types available, and how to use them.

## What is a Model?

A **Model** in the Mindee platform is a configurable and reusable engine designed to process documents. In technical terms, a model is a set of parameters and complex prediction algorithms that perform an inference on files uploaded to Mindee.

Various types of models are available, from simple text recognition to data extraction and classification. Models use Optical Character Recognition (OCR) combined with sophisticated data processing techniques.

Models allow you to automate the process of turning unstructured document images into actionable, structured data. They can be tailored to different document types and business needs, ensuring that only the most relevant information is captured for your workflows.

Each Model contains a dedicated set of tools:

* Configuration: a model is defined by its configuration, that sets the way it should extract data from documents or analyze files. Configuration depends on the model type.
* Settings: overall settings of the model such as processing zone, storage policy, ownership transfer, etc.

## The Models Page

On the Models page, you can view, search, and manage all your models.

Each model is represented as a card showing its name, a preview (if available), and a summary of the main fields it extracts.

All models listed on the Models page are [testable](#live-test) and usable via API.

## Extraction Models

An **Extraction Model** in the Mindee platform is designed to extract structured data from documents using Optical Character Recognition (OCR) combined with data extraction techniques.

More information is available in the section: [Extraction Models](/extraction-models/extraction-models-overview)

## Utility Models

A **Utility Model** in the Mindee platform is designed to perform document analysis, or to preprocess documents in a data extraction workflow.

More information is available for each model type:

* [Split Model Overview](/split-models/split) ⇒ find documents in a multi-page file
* [Crop Model Overview](/crop-models/crop) ⇒ find documents on a single page
* [Classification Model Overview](/classification-models/classification) ⇒ identify file contents
* [Raw Text Model Overview](/raw-text-ocr-models/ocr)⇒ structured text extraction

## Live Test

Once a model is created you can process documents directly on the platform. Use your own files or choose from our samples selection.

You can also view your document history.

This allows easily testing a model before integrating the API and going live to production.

For more information, consult:  [Live Test](/models/live-test).

## Changing Settings

Settings control the high-level options of the model, such as storage duration and processing zone.

For more information on available options, consult:  [Model Settings](/models/model-settings).


# Live Test

The Live Test page from Mindee allows you to test your model, download sample documents or view your document history.

{% @supademo/embed url="<https://app.supademo.com/demo/cmfmnqqnqgjnb39ozfh42qw7z>" demoId="cmfmnqqnqgjnb39ozfh42qw7z" %}

## Overview

The **Live Test** page lets you interactively test your model on real documents and visually inspect the extraction results in one click.

It’s the perfect place to debug, validate, and refine your model’s behavior before moving to production.

## Key Features

* **Side-by-side comparison**: Visually compare the raw document on the left with structured extraction results on the right, ideal for validating model behavior at a glance.
* **Interactive document viewer**: Hover, click, and inspect highlighted fields directly on the image. Tooltips show extracted values and metadata in context.
* **View raw JSON output**: Toggle to display the full JSON response returned by the API, allowing developers to inspect field names, values, coordinates, confidence scores, and nested structures.
* **Instant visual feedback**: See where each field was extracted directly on the document.
* **RAG Toggle**: If your have added documents into your RAG database, you can preview how RAG would influence the output.

## Accessing the Live Test Page

The Live Test is available for all Mindee model types: extraction, crop, split, classification, ....

From your home page on the Mindee platform, simply click on any model.

The first page you see will be the Live Test page.

You can always go back to the Live Test page by clicking on the left-hand model menu.

## Uploading a File for Testing

1. Upload a valid file (PDF or image) directly into the interface.\
   You can click the "Upload a Document" button or drag and drop into the center of the window.
2. Mindee runs the current version of your model on the document.
3. The extracted fields appear on the right-hand side, alongside the image preview.
4. Each field is:
   * **shown on the document** with color-coded bounding boxes overlayed on the document. For Extraction models, this uses the [polygon feature](/extraction-models/optional-features/polygons-bounding-boxes).
   * **shown in the "Extracted Fields" panel** with values, field names, and optionally confidence scores if the [Confidence Score and Accuracy Boost](/extraction-models/optional-features/automation-confidence-score) feature is enabled.
   * also available as raw JSON in the "JSON Response" panel. Note that in Live View, the JSON response always contains the [location information](/extraction-models/sdk-integration/extraction-result#locations) since it is used to highlight the document.
   * interactive — when a field is clicked on the left (document) side, the right (data) side automatically scrolls to the correct section.

<figure><img src="/files/sVkgy7WARnx9XcDuDKca" alt="Overview animation of Mindee Live View"><figcaption><p>By clicking the polygon, you can also see its corresponding field on the right side.</p></figcaption></figure>

### File Limits

<table><thead><tr><th width="170.800048828125">Limit Type</th><th width="152.7999267578125">All Paid Plans</th><th width="162.4000244140625">Free Trial</th></tr></thead><tbody><tr><td>Document File Size</td><td>50 MB</td><td>50 MB</td></tr><tr><td>Number of Pages</td><td>50 pages</td><td>10 pages</td></tr><tr><td>Zip File Size</td><td>100 MB</td><td>100 MB</td></tr></tbody></table>

## Documents History

When testing your model, you'll typically use several documents, and want to use those same documents for testing any changes made to the model.

For this reason you can access the Documents History, where all your tests are stored.

When you access a document in your history after having modified the Data Schema, you'll have a "Outdated Extraction" notice, along with a button to "Rerun Document".

<figure><img src="/files/jIRPhrFxuyEqvsaTiy8b" alt="Live Test - Rerun Document" width="335"><figcaption></figcaption></figure>

### Deleting Documents

You can delete documents and their results at any time.

On the table view, select the documents you wish to delete, then click on the "<i class="fa-trash-can">:trash-can:</i>" icon at the top of the page.

{% hint style="info" %}
Deleting a document is an irreversible action!

The file and its corresponding data will be permanently removed from our system.
{% endhint %}


# Model Settings

Base model settings, model mutability, and legal requirements such as processing zone and document storage.

## Model Name and Cover Image

On the Model Settings page, you'll have the opportunity to modify the model name and the cover image.

Feel free to set them up so that the Model will be easy to find among all your other Models.

## Model ID

Every model has a unique ID which is generated at creation time.

{% hint style="info" %}
Model Templates in the Catalog do not have a Model ID.

The Model ID is generated when the Model is created from the Template. If two users create a Model from the same Template, each Model gets its own unique Model ID.
{% endhint %}

You will need the Model ID in order to [use the Mindee API](/integrations/api-overview).

In the Model Settings page, you can view and copy the Model ID.

The Mindee Support team may request the Model ID to diagnose issues and provide solutions.

## Processing Zone

The **Processing Zone** setting determines the geographic region where your document will be processed. This option can impact compliance with data residency requirements (e.g., GDPR, CCPA).

The processing zone can be set for all models in the [organization settings page](https://app.mindee.com/settings?tab=organization).

If you need to set the processing zone for a specific model, it can be accessed on the model's "General Settings" page.

Available options:

* **No Preference (Default)**\
  Mindee will automatically route your documents to the closest or most available processing region. This is the recommended default for optimal performance and availability, but your data may be processed in Europe and in the United States.
* **Europe**\
  Forces all data processing to occur exclusively within data centers located in Europe (EU). Recommended for organizations subject to GDPR or EU data residency policies.
* **United States**\
  Forces all data processing to occur exclusively within data centers located in the United States. Recommended for organizations subject to U.S. data governance or compliance policies.

{% hint style="info" %}
This is only a zone for *processing*, not extraction or analysis support.

**You can send documents from any country to the Mindee API!**
{% endhint %}

## Storage Policy

This section applies when sending documents via API call.

The **Storage Policy** defines how long extracted data (i.e., the results of document processing) is retained in Mindee’s systems before being permanently deleted. The original uploaded file is never stored to disk.

#### Storage Duration

* Defines the exact number of hours that extracted data is stored.
* **Default:** 12 hours
* **Minimum**: 1 hour
* **Maximum**: 24 hours

Data is automatically deleted as soon as the storage duration expires.

During this time, you may make GET requests to retrieve the payload using its inference ID. After this period, any calls to the inference ID or job ID will result in a 404 error.

This setting helps balance between accessibility of results for your workflow and minimizing retention for privacy and compliance.

#### Delete Extracted Data When Fetched

When **disabled**, data will be retained until the configured **Storage Duration** elapses. **(Default)**

When **enabled**, extracted data is automatically and permanently deleted immediately after:

* the inference is accessed using a GET request (usually when polling)
* the inference is successfully sent to your server via a [webhook](/integrations/webhooks)

The data will be deleted regardless of the Storage Period setting.

Once the Storage Period is passed, the data will be deleted regardless of whether this option is enabled. Once the data are deleted, any calls to the inference ID will result in a 404 error.

This option is recommended for workflows where:

* you only need to access the data once
* you are storing the data from each API call in your system
* there are operational requirements that specify a Zero Data Retention policy

## Copy the Model

It can be useful to copy an existing model for some types of workflows.

You can have a base "template" model that is not called directly, but is used to make derivative models. This way you can have a common base and then modify the Data Schema to account for different providers, geographies, downstream users, etc. Each of these derivative models would be a copy of the "template" model.

You can also us this as way for testing changes to a model. For example you can copy a model used in production, modify the copy, and test the modifications in staging. Once the modifications are tested successfully, switch production over to the new model.

## Lock the Model

To prevent unintended changes once your model configuration is finalized, you have the option to lock the model.

This will prevent all changes to the model's Data Schema and Optional Features.

This is useful for ensuring that the model remains stable for production use.

{% hint style="warning" %}
Locking the Model is an **irreversible action**, you will not be able to unlock it afterwards.

**Note**: you can copy a locked model and modify the copy.
{% endhint %}

## Delete the Model

Models are always active and fully usable.

There is no limit to the number of models you can have.

If you no longer have use of a model, you can delete it at any time.

{% hint style="danger" %}
Deleting a model is an **irreversible action**, all data will be lost forever!
{% endhint %}

Deleting a model also deletes:

* Documents and results stored in the [Documents History](/models/live-test#documents-history)
* Documents and annotations in the [RAG database](/extraction-models/optional-features/improving-accuracy)
* All [Insights](/account-management/insights), including API calls


# Data Processing Policies

Mindee provides flexible options to help you manage how and where your document data is processed, stored, and deleted.

The following features help ensure compliance with your regional regulations and allow control over data retention and privacy.

## Processing Zone

{% hint style="warning" %}
This feature is not available on all [plans](/account-management/plans#feature-comparison).
{% endhint %}

The **Processing Zone** setting determines the geographic region where your document will be processed. This option can impact compliance with data residency requirements (e.g., GDPR, CCPA).

The processing zone can be set for all models in the [organization settings page](https://app.mindee.com/settings?tab=organization).

If you need to set the processing zone for a specific model, it can be accessed on the model's "General Settings" page.

Available options:

* **No Preference (Default)**\
  Mindee will automatically route your documents to the closest or most available processing region. This is the recommended default for optimal performance and availability, but your data may be processed in Europe and in the United States.
* **Europe**\
  Forces all data processing to occur exclusively within data centers located in Europe (EU). Recommended for organizations subject to GDPR or EU data residency policies.
* **United States**\
  Forces all data processing to occur exclusively within data centers located in the United States. Recommended for organizations subject to U.S. data governance or compliance policies.

{% hint style="info" %}
This is only a zone for *processing*, not extraction or analysis support.

**You can send documents from any country to the Mindee API!**
{% endhint %}

## Storage Policy

{% hint style="success" %}
This feature is available on all plans.
{% endhint %}

This section applies when sending documents via API call.

The **Storage Policy** defines how long extracted data (i.e., the results of document processing) is retained in Mindee’s systems before being permanently deleted. The original uploaded file is never stored to disk.

#### Storage Duration

* Defines the exact number of hours that extracted data is stored.
* **Default:** 12 hours
* **Minimum**: 1 hour
* **Maximum**: 24 hours

Data is automatically deleted as soon as the storage duration expires.

During this time, you may make GET requests to retrieve the payload using its inference ID. After this period, any calls to the inference ID or job ID will result in a 404 error.

This setting helps balance between accessibility of results for your workflow and minimizing retention for privacy and compliance.

#### Delete Extracted Data When Fetched

When **disabled**, data will be retained until the configured **Storage Duration** elapses. **(Default)**

When **enabled**, extracted data is automatically and permanently deleted immediately after:

* the inference is accessed using a GET request (usually when polling)
* the inference is successfully sent to your server via a [webhook](/integrations/webhooks)

The data will be deleted regardless of the Storage Period setting.

Once the Storage Period is passed, the data will be deleted regardless of whether this option is enabled. Once the data are deleted, any calls to the inference ID will result in a 404 error.

This option is recommended for workflows where:

* you only need to access the data once
* you are storing the data from each API call in your system
* there are operational requirements that specify a Zero Data Retention policy

## Best Practices

* **For testing and setup:** Use a longer Storage Duration (up to 24 hours) to allow time to check and validate processed results. This is particularly useful when setting up a webhook workflow.
* **For general production use:** Set a relatively short Storage Duration (3-5 hours).
* **For privacy or Zero Data Retention:** Always enable **Delete When Fetched.** Best to also set a short Storage Duration (1 hour) in case of failed calls (i.e. network failure on GET).
* **For compliance:** Select the Processing Zone that matches your legal and regulatory requirements.

### Example Configuration

* **Processing Zone**: Europe
* **Storage Duration**: 4 hours
* **Delete When Fetched**: Enabled

With this setup, all documents are processed in EU data centers, results are stored for a maximum of 4 hours, and deleted automatically once you first download them.

## **Configure the Data Processing Policy**

{% hint style="warning" %}
This feature is not available on all [plans](/account-management/plans#feature-comparison).
{% endhint %}

{% @supademo/embed url="<https://app.supademo.com/demo/cmevfnfze75icv9kqgqida892>" demoId="cmevfnfze75icv9kqgqida892" %}

## Frequently Asked Questions

<details>

<summary><strong>Does Mindee use my documents for training?</strong></summary>

No, all documents are the property of their organization and are not used for training Mindee models.

We only use your documents internally to help us resolve a ticket you opened with our support team. Since we don't have access to documents sent by API, we would ask you explicitly for samples during the support request investigation, if needed.

</details>

<details>

<summary><strong>Are my documents or their data shared or sold to 3rd parties?</strong></summary>

No, never.

Documents and their data are only accessible to the users of the organization that the model belongs to.

Mindee never sells documents or their data to 3rd parties.

</details>

<details>

<summary><strong>Is Mindee GDPR &#x26; SOC 2 Type II compliant?</strong></summary>

Yes. Mindee V2 is fully GDPR-compliant and maintains its SOC 2 Type II certification, ensuring rigorous security controls, data protection standards, and transparency for enterprise needs.

</details>

<details>

<summary><strong>Can I request data deletion at any time?</strong></summary>

Yes. You can request full or partial data deletion at any time by using the deletion features available in your [model settings](/models/model-settings#delete-the-model), [organization settings](/account-management/organizations#delete-organization), or [account settings](/account-management/account-settings#delete-account).

Mindee ensures that your data is securely and permanently removed upon request.

You can also contact our support team if you are unsure.

</details>

<details>

<summary><strong>Can you sign a BAA for HIPAA compliance?</strong></summary>

We can, however these agreements are limited to enterprise users.

If you have an enterprise plan, contact your dedicated account manager for more information.

</details>


# Credit Cost

Information on credit consumption for processing documents using a given model.

Each file successfully processed by a model will consume credits based on the number of pages contained in the file.

Files that are sent but not processed (i.e. job status "Failed", or HTTP 400) do not consume credits.

The model's "Credit Cost" page shows the exact amount of credits consumed per page processed.

For files having more than one page, such as PDF or TIFF, this is every page of the file.\
Single image files count as one page.

For example, assuming 1 credit per page:

* Single page text PDF file ⇒ 1 credit
* 5 page text and image PDF file ⇒ 5 credits
* A JPEG file ⇒ 1 credit

{% hint style="info" %}
Ratio of 1 credit per page is for demonstrative purposes.

Depending on the model type and options activated, credits consumed per page may be lower or higher.

**Always check the model's Credit Cost page for the exact amount of credits consumed.**
{% endhint %}


# Extraction Model Overview

Extraction models enable structured data extraction based on fields, which can be configured for any document type.

## What is an Extraction Model?

An **Extraction Model** in the Mindee platform is a type of [model](/models/models-overview) designed to extract structured data from documents using Optical Character Recognition (OCR) combined with advanced data extraction techniques.

Each model defines a set of fields called a [Data Schema Overview](/extraction-models/data-schema). Fields could include "Supplier Name," "Invoice Number," or "Total Amount" for an Invoice Model, as an example. The system will then identify and extract data from uploaded files according to these field definitions.

You can create extraction models for any type of document including: invoices, receipts, passports, ID cards, financial statements, etc.

Models allow you to automate the process of turning unstructured document images into actionable, structured data. They can be tailored to different document types and business needs, ensuring that only relevant information is captured for your workflows.

Each Extraction Model contains these dedicated tools and features:

* [Data Schema Overview](/extraction-models/data-schema): a model is also defined by a Data Schema, that sets the list of fields the API should extract for a given type of documents.
* [Continuous Learning (RAG)](/extraction-models/optional-features/improving-accuracy): it allows the user to give additional instruction on some examples with unexpected behavior to durably improve the extraction performance of the model.
* [Model Settings](/models/model-settings): overall settings of the model such as processing zone, storage policy, ownership transfer, etc.

## Create an Extraction Model <a href="#creating-from-scratch" id="creating-from-scratch"></a>

From the Mindee Platform main page, click on **"Create your document AI model"**.

### Create From the Model Catalog

Chances are the document you're trying to automate is a known type of document, and likely to be already present in our Model Catalog: Invoice, Receipt, Passport, Financial Document, ID Card...\
\
These are model templates with Data Schemas created by our Data Science team, allowing you to get started quickly by using a set of predefined fields.

You can search our catalog for suitable templates, then simply click on the one that best fits your needs. This will create a new model in your organization's account with its own model ID, allowing you to use and modify it.

{% hint style="info" %}
**Catalog Models come with generic extraction fields.**

You'll likely need to add and/or modify fields to fit your needs and business case.
{% endhint %}

### Create a Custom Model <a href="#creating-from-scratch" id="creating-from-scratch"></a>

If none of the catalog models fit your needs, you can start from scratch.\
When creating a new Model, click on **"Custom document"**.

Next, choose how to create the custom model:

* Describe as precisely as possible the document(s) the model will process.\
  Best when documents processed show variance.

or

* Upload a sample file representative of documents to process.\
  Best when documents processed are similar.

Our AI Agent will help you to quickly define an initial data schema that you'll be able to adjust later.

This step will also generate the model's unique ID.

{% hint style="info" icon="lightbulb" %}
**We recommend creating an initial model quickly.**

Once created, refine the new model's [Data Schema Overview](/extraction-models/data-schema) on the platform, using the [Live Test](/models/live-test) feature to fine-tune your fields and guidelines.
{% endhint %}

## Modifying Your Extraction Data Schema

Once you have created a model, you'll likely want to modify its [Data Schema](/extraction-models/data-schema) to better suit your needs.

Even if the model was created from a model template in the Catalog, it is unique to your account and fully customizable.

Navigate to the Data Schema page where you can adjust fields, update configurations, and customize settings according to your requirements.

### Using the AI Assistant

The Mindee AI Assistant is a powerful tool that can help you get the most of your model. It is the preferred method of modifying the Data Schema.

The Assistant is available in a dialog box in the model's Data Schema page.

Be as accurate as possible with field names, exact names and values should be in quotation marks.\
\
Example sentences to modify your data schema with the AI Assistant:

* *Add a new field: "document id"*\
  ⇒ Will create a new text field.
* *Add a new field: "is past due"*\
  ⇒ Will create a new boolean field\
  Hint: boolean field names should start with "is" or "has".
* *Rename the "date" field to "invoice date"*
* *Change the "document\_type" field to a classification field. The expected classes are "INVOICE" and "RECEIPT"*

## Optimizing Model Results

In order to optimize the accuracy of a given model, the first step is to [fine-tune the Data Schema](/extraction-models/data-schema-best-practices).

In particular, start by asking the Mindee AI Assistant to [automatically optimize](/extraction-models/data-schema-best-practices#automated-optimization) your Data Schema.

If you started from a template in the catalog, it is usually necessary to adapt it to your specific documents and use case.

After this step, if there are still some lingering issues or unexpected behavior, there are further refinements available.

### Problems With Specific Templates

When the overall accuracy is good, but there are problems on specific document templates.

For this you'll want to activate [Continuous Learning (RAG)](/extraction-models/optional-features/improving-accuracy). This allows adding guidance to specific templates, meaning you can target problem documents while leaving others unaffected.

### High Accuracy Required

When you need very high precision for all the documents that you process.

Consider enabling [Confidence Score and Accuracy Boost](/extraction-models/optional-features/automation-confidence-score). This uses multiple models for higher accuracy **and** flags problematic fields.

Flagged fields can be sent for [human review](#end-user-review) or the document can be rejected entirely, based on business rules defined on your side.

### End User Review

When you are displaying the documents to end users, typically when they can review and correct the extracted data.

You can enable [Polygons (Bounding Boxes)](/extraction-models/optional-features/polygons-bounding-boxes) so that the location of the extracted fields can be shown (this is always activated in the [Live Test](/models/live-test)). In this way, it will be much easier for your users to find erroneous fields correct them.

This is even more powerful when combined with [confidence scores](/extraction-models/optional-features/automation-confidence-score), you can flag fields needing attention directly in your forms.

### Combined Benefits

These features work together to create a system that becomes more accurate as you use it, reducing manual corrections and improving automation success rates.


# Data Schema Overview

General description of an extraction model's Data Schema.

An **Extraction Data Schema** defines which data should be extracted, and in what way, from documents sent to the model.

The Data Schema is the foundation of your model. All other features will use and depend on the Data Schema.

You can think of the Data Schema as a map to the documents you send to the model for extraction.

This map includes which data to extract, how to format the data, pitfalls to avoid, etc.\
\
To do this, the Data Schema is composed of Fields (or data points), each field having its own configuration.

{% hint style="info" icon="lightbulb" %}
Use the [Live Test](/models/live-test) when working on your Data Schema to quickly validate changes.
{% endhint %}

## Fields Overview

A Data Schema is primarily composed of fields.

A field describes a single data point to extract in the document.

Each field has the following properties:

* Title - human-readable
* Name - machine-readable, used as the field's key in the API return
* Type - the type of data to extract, details in: [#field-types](#field-types "mention")
* Description (optional) - provides extra context on how the field is used
* Guidelines (optional) - provides instructions to better extract the field

## Field Types

A field's type determines how it will be formatted when returned by the API.

The type also gives an indication on what to look for in the input file.

### Base Types

<table><thead><tr><th width="191.2000732421875">Field Type</th><th>Description</th></tr></thead><tbody><tr><td>Text</td><td>A sequence of characters representing textual data, a string value.</td></tr><tr><td>Number</td><td>Numeric data which can be a whole value (integer) or a decimal value (floating point).</td></tr><tr><td>Date</td><td>A specific year, month, and day, formatted as a <code>YYYY-MM-DD</code> date or <code>YYYY-MM-DD HH:mm:ss</code> for date and time.</td></tr><tr><td>Classification</td><td>A defined list of categories or types to match. Category names are text strings.</td></tr><tr><td>Boolean</td><td><p>Represents two possible values: <code>true</code> or <code>false</code> </p><p>Should be used for checkboxes.</p><p>Hint: boolean field names should start with "is" or "has".</p></td></tr><tr><td>Nested Object</td><td><p>A complex data type containing multiple subfields or properties.</p><p>Used to group related elements, such as the components of an address.</p><p>Subfields may contain a single value or a list of values.</p><p>Only one level of nesting: subfields may not be nested objects.</p></td></tr><tr><td>Object Detection</td><td>Detect the location of a document feature, such as a logo, signature, photo, etc.</td></tr><tr><td>Barcode</td><td><p>Detect the location of a 1D barcode (i.e. UPC, EAN) or a 2D barcode (i.e. QR Code, Data Matrix, 2D-DOC, PDF417).</p><p>Additionally, attempt to decode the contents of the barcode as a string value.</p></td></tr></tbody></table>

{% hint style="info" %}
All fields can have an empty (null) value.
{% endhint %}

### **Array Types**

Any field type can be made into an array, a list of values.

Simply enable "Multiple items can be extracted" when creating or modifying the field.

The return type will be an array of the base type, for example a list of text values, or a list of numbers.

It is possible to have a list of nested objects, but not a list of lists.

{% hint style="info" icon="lightbulb" %}
In some cases, there can be duplicate items, for example when the same value appears on several pages.

Enable "Filter out duplicates from the list of items" to fix this.
{% endhint %}

## **Field Examples**

Some examples for the best field types to use, given a basic invoice extraction Data Schema.

<table><thead><tr><th>Field Name</th><th width="203.5">Field Type</th><th>Example Return Value</th></tr></thead><tbody><tr><td><strong>Supplier Name</strong></td><td>String</td><td><code>Acme Supplies Ltd.</code></td></tr><tr><td><strong>Supplier Logo</strong></td><td>Object Detection</td><td>Polygon around the logo</td></tr><tr><td><strong>Supplier Company Registration</strong></td><td>Nested Object</td><td><em>See sub-fields below</em></td></tr><tr><td><em>Supplier Company Registration.</em><strong>Number</strong></td><td>String</td><td><code>CRN-20250123</code></td></tr><tr><td><em>Supplier Company Registration.</em><strong>Type</strong></td><td>Classification</td><td><code>VAT NUMBER</code></td></tr><tr><td><strong>Invoice Date</strong></td><td>Date</td><td><code>2025-06-10</code></td></tr><tr><td><strong>Is Past Due</strong></td><td>Boolean</td><td><code>false</code></td></tr><tr><td><strong>Total Amount</strong></td><td>Number</td><td><code>1540.75</code></td></tr><tr><td><strong>Taxes</strong></td><td>Nested Objects Array</td><td><em>See sub-fields below</em></td></tr><tr><td><em>Taxes[0].</em><strong>Rate</strong></td><td>number</td><td><code>0.185</code></td></tr><tr><td><em>Taxes[0].</em><strong>Base</strong></td><td>number</td><td><code>1300.00</code></td></tr><tr><td><em>Taxes[0].</em><strong>Amount</strong></td><td>number</td><td><code>240.75</code></td></tr></tbody></table>

## Overall Guidelines

In addition to individual field guidelines, an overall (or global) guideline can be used in your Data Schema.

The overall guideline text will apply to all or some fields, depending on your instructions.

Use overall guidelines when you want to:

* Generalize instructions to several specific fields. For example:
  * *"Number fields related to amounts should always have 3 decimal places."*
  * *"Country fields should return the ISO alpha-3 code of the country."*
* Provide general instructions or context for all fields. For example:
  * *"Ensure ASCII compliance by removing all diacritics from return values."*

{% hint style="info" %}
You may put any number of unrelated guidelines in the text, for example all of the samples above.

For best results, separate each different guideline with a new line.
{% endhint %}

## Document Text

### Language

You can specify a field's *Title*, *Name*, *Description*, and *Guidelines* in most languages.

This also applies to the Data Schema's [#global-guidelines](#global-guidelines "mention").

Mindee models can process documents in most languages.

This includes, but is **not limited** to:

* European languages: English, French, Spanish, German, Italian, Portuguese, Russian, Greek, etc
* Asian languages: Hindi, Bengali, Turkish, Urdu, Farsi, Armenian, etc
* East Asian languages: Japanese, Mandarin, Korean, Vietnamese, etc
* Semitic languages: Arabic, Hebrew, Amharic, etc
* African languages: Swahili, Yoruba, Zulu, etc

**Note:** while the models can understand these languages, we are not able to provide in-depth support for all languages.

### Handwriting

Mindee models are able to recognize (or OCR) handwritten texts. Accuracy for handwriting is on average a bit less than printed text.

## Technical Limitations

### Number of Fields in the Data Schema

The recommended maximum number of fields is 25 for a Data Schema.

While there will be no errors, beyond this number response times will increase.

### Names of Fields

The field *name* must only contain:

* lowercase Latin letters without accents (a-z)
* numbers (0-9)
* underscores (`_`), but neither first nor last characters can be an underscore.

## Next Steps

Now that you're familiar with the different components of the Data Schema, you'll want to take a look at tips for [building an accurate Data Schema](/extraction-models/data-schema-best-practices).


# Data Schema Best Practices

Best practices for designing a Data Schema that improves extraction accuracy.

{% hint style="info" %}
We recommend taking a look at [Data Schema Overview](/extraction-models/data-schema) first.\
This will help you understand the terms used on this page.
{% endhint %}

Mindee models are not "trained" using manually annotated documents, instead the Data Schema is adjusted.

This means that to get the best possible extraction data from a model, the key is to ensure that the Data Schema is clear and optimized.

Any additional features you activate to increase accuracy, such as [Continuous Learning (RAG)](/extraction-models/optional-features/improving-accuracy) or [Confidence Score and Accuracy Boost](/extraction-models/optional-features/automation-confidence-score) will rely heavily on the Data Schema.

## **Automated Optimization**

Before modifying the Data Schema manually, first start by running our provided AI-assisted optimization tool.

The tool will analyze your documents and determine how well the Data Schema fits them. From there it will provide suggestions for improving the accuracy of the extractions.

The process is completely automated, you just need to provide some sample documents.

### Running the Analysis

To start the analysis, first upload at least 5 documents using the [Live Test](/models/live-test).

Next, go to the Data Schema page and click the "<i class="fa-message">:message:</i> AI Assistant" button:

<figure><img src="/files/iDWUAZPGoFdODJ2gQLgV" alt="AI Assistant Button" width="158"><figcaption></figcaption></figure>

In the dialog box, click the "<i class="fa-wand-magic-sparkles">:wand-magic-sparkles:</i>" (magic wand) button:

<figure><img src="/files/cokgCAIRrk4FYocTLA9P" alt="AI Assistant Auto Optimize Button" width="408"><figcaption></figcaption></figure>

Now you can sit back, grab a cup of tea, and wait for the analysis to complete.

The model's Data Schema will be temporarily locked during this process. You can safely navigate away from the model page, the analysis will keep running in the background.

### Applying the Analysis Results

Once the analysis is complete, the chat window will fill with all of the suggested improvements.

For each field, you'll see:

* a checkbox
* the current "Baseline" description and guidelines, in red
* the proposed "Optimized" changes, in green

To apply the proposed changes, check each field for which you want to apply the changes.

When you have selected all the fields you wish to modify, click the "Apply Selected" button:

<figure><img src="/files/YheWPxiPKsm8qHgvWplB" alt="AI Assistant Optimize Apply Changes Button" width="389"><figcaption></figcaption></figure>

You'll receive a confirmation message in the dialog window once the changes have been applied.

### Testing the Optimized Data Schema

To test and validate the changes, head on over to the Live Test, then click on the "Documents History" tab.

From there, open the documents and click on the "Rerun Document" button (it's a restart icon):

<figure><img src="/files/jIRPhrFxuyEqvsaTiy8b" alt="Rerun the Document Button" width="335"><figcaption></figcaption></figure>

## **Best Practices for Fields**

The heart of the Data Schema. The various properties of fields all have a role to play in getting the best possible accuracy.

### **Field Name and Title**

The field *Name* is automatically generated from the field *Title*. You can also modify the *Title* afterwards.

Both the *Name* and *Title* are used during processing (inference).

Use clear, simple names that will precisely describe the field you want to extract.\
The goal is to avoid any possible confusion between data points present in the document.

As an example, let's extract the name of the company that issued an invoice.

In our Data Schema, we've used the field *Name*: `supplier_name`\
It clearly tells the model to extract only the name of the invoice supplier.

:white\_check\_mark: you could also use `vendor_name`, it has a similar meaning and equivalent precision.

:warning: `supplier` might work but is too broad: which information about the supplier is needed, exactly?

:warning: `company_name` might work but is ambiguous: we know you need the name of the company, but we don't know if company stands for supplier or customer.

:x: `company` will likely not work as expected: we know neither what information you need nor which company is concerned.

### Field Type

Try to use one of the [field types](/extraction-models/data-schema#field-types) that best matches how the field is used and how it appears on the document.

For example, while you could use a string for `due_date`, a date field type is definitely better.

### Field Description

The field's *Description* has an impact on the model's performance.

Use it to describe what the field represents, and/or of what its use is to you.

For example, the `supplier_name` field could have:

> The name of the supplier.
>
> Used in internal processing to match our supplier ID with the name found on the document.

### Field Extraction Guidelines

Sometimes changing the field name and type is not enough to explain what you need for one field.\
In that case you can add extraction *Guidelines* to the field.

Use natural language to explain how to properly extract the data, and/or any extra steps like formatting.

For instance, with `supplier_phone_number`, adding the following extraction guidelines could be useful:

> If you find several phone numbers in the document, use the phone number of the supplier headquarters.
>
> Always reformat the data to match the international phone number format, as follows: +1-212-867-5309

## Relative Importance of Field Properties

Not all field properties have the same importance or weight when it comes to how the models process files.

Additionally, not all types of fields are handled the same way.

In the following table, "Normal Fields" are those that extract textual information from the document (text, dates, numbers, etc), whether they are simple fields, lists, or nested object fields.

"Object Detection" refers to specific processing to extract polygons of various elements on the document, such as signatures, ID photos, etc.

<table><thead><tr><th width="208.0001220703125">Property</th><th width="261.4000244140625">Normal Field Usage</th><th>Object Detection Usage</th></tr></thead><tbody><tr><td>Name</td><td><strong>Most important</strong></td><td>Not used</td></tr><tr><td>Title</td><td>Important</td><td><strong>Most important</strong></td></tr><tr><td>Description</td><td>Complementary</td><td>Not used</td></tr><tr><td>Guidelines</td><td>Complementary</td><td>Not used</td></tr><tr><td>Classification Values</td><td>Very important (only for classification fields)</td><td>Not used</td></tr></tbody></table>

## Less Is More

It can be tempting to give very detailed instructions in guidelines and descriptions. However, in many cases this is actually counterproductive and will lead to diminished accuracy.

Here is an example of too many details (don't do this):

> The order number is usually next to the words "order no" on the invoice, but sometimes there is no order number, so it will be next to the words "customer invoice no". It is usually present on the first page, in a green box.

What's wrong here? Well, the first thing to understand is that you are not giving instructions to a human, but to a machine. Machines prefer concise instructions. On the other hand, this machine has been trained on millions of documents, and is capable of determining the location of a value field on its own.

So what's left is just the instruction about the *customer invoice* and *order number*.

A simpler, better version would be:

> Use the value of "customer invoice number" if "order number" is missing in the document.

## Remove Ambiguity with Extra Fields

In some cases, it can be beneficial to add extra fields you don't actually need in order to remove ambiguity on the data you do need.

Let's say you are processing invoices, and need to extract the "Reference Number". When processing the invoices, you notice that, sometimes, the "Order Number" is picked up as the reference number.

A first logical step would be to add a guideline, something like *"NEVER use the 'Order Number' to populate this field"*. While this should work in **most** cases, the distinction between a Reference and Order may not be perfectly clear to the model.

Adding more text or more detailed information is potentially [counterproductive](#less-is-more).

A potential fix would be to add the "Reference Number" field in your Data Schema, in addition to the "Order Number" field. This way the ambiguity is lifted, it is now very clear to the model that these are separate data points.

Then, in your data processing flow, simply ignore the extra field.

As a reminder, the number of fields in the Data Schema has no impact on pricing.

## Best Practices for Different Regions or Languages

It's important to first make the distinction between *representation* and *content*.

Representation meaning that equivalent content can be showed or displayed differently (for language, this is a translation).

Content meaning that the structure of the data is different, regardless of its representation (language).

In the context of a Data Schema, the optimal configuration mainly depends on whether the data you need to extract (the content) changes depending on the region or language.

Typically the follow-up question is: should I use a single model or multiple models?

### When to use different Models

If the data changes considerably, it will be beneficial to have different models for different regions.\
For example:

* different tax lines/calculations on invoices
* specific fields on ID documents
* different reporting on energy bills

Having different models to better fit the data to extract will generally provide more accurate results.\
It will also allow having specific [#field-extraction-guidelines](#field-extraction-guidelines "mention").

### When to use a single Model

On the other hand, if the data to extract does not vary significantly, even when the language changes, there is generally no need to have different models.\
For example:

* same document, different language (multilingual countries like Belgium, Canada, India, etc)
* same data to extract, even if the document changes (only a subset of data is required)

## Frequently Asked Questions

<details>

<summary><strong>When should I use Data Schema guidelines vs RAG guidelines?</strong></summary>

Both work in the same way to provide extra context for the model to better identify data and extract them.

The major difference is *when* they are applied:

If the guideline is to be applied to all documents ⇒ use the Data Schema guideline.

If the guideline only applies to a specific template ⇒ use the RAG guideline.

</details>

<details>

<summary><strong>Can the model correctly interpret positions like top, bottom, etc?</strong></summary>

Mindee models are multimodal, meaning they use both textual and visual information.

As such it is possible to write guidelines conveying positional information, for example:

> The supplier logo will always be at the top of the page.

</details>

<details>

<summary><strong>Can I reference a specific location or text on the page?</strong></summary>

Each page of the document is processed as a whole, and the models have both a visual component and a textual component (multimodal models).

It is therefore possible to reference other locations and texts of the page in guidelines.

This is especially useful to remove ambiguity when similar data is found in multiple locations on the page.

Referencing another location:

> The correct phone number is below the customer name.

Specifying a location and text:

> The correct customer name is found in the first line of the ship-to address.

</details>

<details>

<summary><strong>Does the model work better with field information in English?</strong></summary>

In our own testing, major European languages such as French, Spanish, or German, are comparable to English in terms of modal accuracy.

In some cases, it *may* be beneficial to define the Data Schema, such as field names and descriptions, using the language present in the document. This can help improve accuracy, in particular when terms not easily translatable are used in the field definition.

For example, a model for French invoices may use "TVA Intracommunautaire" as the field name rather than "Intra-Community VAT", although both will work.

Best is to try both using the [Live Test](/models/live-test) feature.

</details>

## Next Steps

Once you are satisfied with the results of your Data Schema, you'll want to [connect your platform](/integrations/api-overview) and start processing documents.


# Optional Features

Optional model features to enhance document processing.

## Overview

Models can have optional processing features activated depending on your requirements.

{% hint style="info" %}
Not all features are available on all plans.

Check the [Plans and Credits](/account-management/plans#feature-comparison) section for more information.
{% endhint %}

There are two ways of enabling and disabling features: on the Platform, or via API calls.

### Set Default Activation on the Platform

When setting the activation state of a feature on the Platform, this will be the default.\
All API calls will use the default state unless explicitly set otherwise during the API call.

This is useful for project managers, as it allows activating or deactivating a feature across all API calls on that model.

Anyone with write access to the model can set the option's default value.

### Set Activation During API Call

Default activation states may be overridden on a per-call basis via the API.

This is useful for developers for a variety of different scenarios, for example:

* local testing, comparing results on the same files with the feature active and inactive
* conditional activation of a feature based on business rules: document template, file origin, etc
* conditional activation of a feature based on your end-user's configuration
* etc ...

Anyone with access to the API (via an API Key) can enable or disable a feature when making an API call.

Details on setting features in the API: [Client Configuration](/integrations/client-libraries-sdk/configure-the-client#optional-features-configuration)

## Available Extraction Features

* [Raw Text (Full OCR)](/extraction-models/optional-features/raw-text-full-ocr)\
  Add the full text content of your documents to the API response.
* [Continuous Learning (RAG)](/extraction-models/optional-features/improving-accuracy)\
  Enhance extraction accuracy with Retrieval-Augmented Generation using your own documents.
* [Polygons (Bounding Boxes)](/extraction-models/optional-features/polygons-bounding-boxes)\
  Add the polygon coordinates of each extracted field to the API response.
* [Confidence Score and Accuracy Boost](/extraction-models/optional-features/automation-confidence-score)\
  🚀 Boost the precision and accuracy of all extractions.\
  Add a confidence score to each extracted field.


# Continuous Learning (RAG)

Enhance extraction accuracy with Retrieval-Augmented Generation (RAG) using your own documents.

## Overview

When using an Extraction Model, sometimes the extraction accuracy is not satisfying on a given document or a particular template.

If you have already followed the [Data Schema Best Practices](/extraction-models/data-schema-best-practices), and results are still not satisfactory, all hope is not lost.

You have a way to durably improve the performance of the model for the next predictions you'll make. The solution you need is the RAG feature.

{% embed url="<http://app.supademo.com/demo/cmfb3c7166gdc39ozj39wogrz>" %}

## Basics of RAG Inner Workings

Retrieval-Augmented Generation or RAG, is an approach in artificial intelligence that combines the strengths of retrieval-based models with generative LLM text production.

Essentially, RAG leverages a rich database of source documents and embeddings to augment the knowledge base from which it draws responses. This means that the system retrieves specific and context-based data snippets and then uses them to generate semantic and fine-tuned responses that are both creative and factually accurate.

### Overall Steps of RAG

1. Annotating a given sample, which means giving the model the difference between the extracted data and the expected value(s). When activated, this example will be added to the RAG Database for the future.
2. Retrieval: When a next document will be send with the RAG option activated, the model will try to search for a similar example in the existing database. The question here is : "Maybe there is an existing example where I need to follow the instructions so that I'm not doing a same mistake again". If no example found, no need to augment the prediction. If an examples is matched in the RAG Database, here comes step 3.
3. Augmented Generation: A document was matched in the RAG Database. The model will use the instructions you gave on the RAG sample to make a better prediction this time. The prediction generated is augmented with an existing context helping the model to be better this time.

<figure><img src="/files/9YLUqele5RJX9NRC7RbA" alt="RAG process overview" width="563"><figcaption></figcaption></figure>

## Set Up the RAG Database

To use RAG on a given Extraction Model, you'll need to first set up the RAG database which will contain your documents and data.

You only need a single document to get started.

Access the database by clicking on the "Continuous Learning (RAG)" link in the model's configuration section.\
\
You can then enrich your RAG database by uploading documents with problematic responses.

You need to annotate the document, ticking the fields you want to be covered by the RAG augmentation on this template. You can also add additional guidelines using plain language.

{% hint style="success" %}
Most of the time, the annotation is sufficient to make the model understand the issue.\
We recommend using the guideline only when the annotation doesn't solve the problem.
{% endhint %}

Once this document is annotated, **be sure to validate it**, and go to the Live test tab.\
\
You should upload a document, and leave "Show RAG extraction" ticked.

Ideally, pick a document with the same template (another invoice from the same supplier for instance), but not exactly the one you used in the RAG database. You will see the before/after predictions and should be able to check that the extra instructions were taken into account to augment properly the prediction.\
\
In the future, the documents respecting the same template should be augmented, which should increase a lot the performances on this given template. For other types of documents, the behavior remain the same, which means that RAG is improving the result with no regression on other documents.

To see which documents are being used, in the database the columns "Matched" and "Last Matched" indicate how many times the document was used and the last time it was used, respectively.

## Activate RAG

For best results, be sure to follow the [Data Schema Best Practices](/extraction-models/data-schema-best-practices) before activating this feature.

If the Data Schema is not optimized, this feature may not significantly improve model accuracy.

### Activate RAG on the Platform

When setting the activation state of a feature on the Platform, this will be the default.\
All API calls will use the default state unless explicitly set otherwise during the API call.

This is useful for project managers, as it allows activating or deactivating a feature across all API calls on that model.

Anyone with write access to the model can set the option's default value.

Click on "Data Schema" then "Optional Features" tab. There you can activate the "RAG Processing" feature.

### Activate RAG via API Calls

If you need finer-grained control over when the feature is used, you can activate or deactivate it when making API calls.

This allows dynamically setting the activation state using your internal business or domain logic.

Check the [Client Configuration](/integrations/client-libraries-sdk/configure-the-client#optional-features-configuration) section if using our [Client Libraries / SDKs](/integrations/client-libraries-sdk).

## Frequently Asked Questions

<details>

<summary><strong>How is my data protected during RAG processing?</strong></summary>

Your data remains private and isolated within your model.

All documents and annotations added to your RAG database are encrypted in transit.

</details>

<details>

<summary><strong>Is my RAG data shared with other users or sold to 3rd parties?</strong></summary>

No, never.

RAG data is only accessible to the users of the organization that the model belongs to.

Mindee never sells data to 3rd parties.

All storage and retrieval of RAG data is done on Mindee's dedicated servers.

</details>

<details>

<summary><strong>Does Mindee use my RAG data for training?</strong></summary>

No, all documents are the property of their organization and are not used for training Mindee models.

We only use RAG data internally with your explicit knowledge and prior consent. This is limited to using your documents to help resolve a ticket you opened with our support team.

</details>

<details>

<summary><strong>What happens to my RAG database if I downgrade my plan?</strong></summary>

We won't delete any of the files you've uploaded into your RAG document database if you downgrade, however you will not be able to modify or validate them.

The guidelines you have given to your RAG documents will be taken into account in your inferences with the lower plan, even if that is above the plan's features.

</details>


# Confidence Score and Accuracy Boost

Enable automated workflows by enhancing model accuracy and measuring field confidence.

## Overview

The **Automation** feature in Mindee's platform represents a major step forward in enhancing both the **accuracy** and **reliability** of document data extraction. Designed to support robust and scalable automation workflows, this feature is built on two core capabilities:

1. **Enhanced accuracy** using model ensemble algorithms
2. **Confidence scoring** for all types of extracted fields

Automation aims to solve two common challenges in intelligent document processing:

* **Maximizing extraction quality** in the face of variable and noisy document formats
* **Providing actionable trust signals** so that systems can handle uncertain extractions appropriately

By combining multiple models and analyzing their agreement, Automation ensures that the most reliable prediction is selected for each field, while transparently communicating how confident the system is in that prediction.

## Use Cases

### Full Automation

By leveraging **confidence score thresholds**, you can selectively **automate decisions** in your processing pipeline, triggering downstream actions only when extractions meet a predefined reliability level.

For example, fields marked with a `High` or `Certain` confidence score can be automatically approved and pushed to your ERP or CRM system, while extractions with `Low` or `Medium` confidence can be routed for human review or fallback logic. This selective gating mechanism allows teams to implement **fully automated flows** for clean, predictable documents, while still handling edge cases gracefully.

#### **Use cases examples**:

* Auto-validating invoice totals and tax fields before ingestion into an accounting system
* Auto-approving identity document extractions for KYC when confidence is high
* Automatically flagging low-confidence vendor names or dates for manual verification

### Efficient human validation

To make confidence levels easily understood by end-users, each confidence score returned by Automation can be associated with a **color-coded indicator**.\
This visual feedback is especially useful in UI-driven workflows, where operators need to scan, validate, or correct extractions quickly.

The default color scheme is as follows:

| Confidence Level | Label     | Color Code | Suggested Action       | Description                                                              |
| ---------------- | --------- | ---------- | ---------------------- | ------------------------------------------------------------------------ |
| 🟦 Certain       | `Certain` | Blue       | Safe for automation    | Full confidence, human-level precision                                   |
| 🟩 High          | `High`    | Green      | Can be auto-processed  | Model consensus is strong; prediction is likely accurate.                |
| 🟧 Medium        | `Medium`  | Orange     | Optional review        | Some confidence, but context or format may impact correctness.           |
| 🟥 Low           | `Low`     | Red        | Manual review required | Extraction is uncertain or likely incorrect. Model disagreement is high. |

This color-coding system allows product teams to **highlight uncertainty directly in the user interface**, enabling faster decisions, reducing cognitive load, and streamlining exception handling.

## Activate Confidence Scores

For best results, be sure to follow the [Data Schema Best Practices](/extraction-models/data-schema-best-practices) before activating this feature.

If the Data Schema is not optimized, this feature may not significantly improve model accuracy.

### Activate Confidence Scores on the Platform

When setting the activation state of a feature on the Platform, this will be the default.\
All API calls will use the default state unless explicitly set otherwise during the API call.

This is useful for project managers, as it allows activating or deactivating a feature across all API calls on that model.

Anyone with write access to the model can set the option's default value.

{% @supademo/embed url="<https://app.supademo.com/demo/cmeie3irw9fe7h3pytuktflxs>" demoId="cmeie3irw9fe7h3pytuktflxs" %}

### Activate Confidence Scores via API Calls

If you need finer-grained control over when the feature is used, you can activate or deactivate it when making API calls.

This allows dynamically setting the activation state using your internal business or domain logic.

Check the [Client Configuration](/integrations/client-libraries-sdk/configure-the-client#optional-features-configuration) section if using our [Client Libraries / SDKs](/integrations/client-libraries-sdk).

{% hint style="info" %}
When the **Automation** feature is not activated, the `confidence` attribute in the response will always be `null`.
{% endhint %}

### Using Confidence Scores in Processing

You can easily add various business and/or processing logic rules in your code to handle complex workflows.

Take a look at the [Extraction Result](/extraction-models/sdk-integration/extraction-result#confidence) section for implementation details.

## Towards 100% Automation

By combining confidence-based automation with Mindee’s **RAG-powered continuous learning loop**, you can drive your workflows toward **near 100% automation**.

Low-confidence extractions are not only flagged for human validation, but also used as feedback signals to refine models dynamically, through retrieval-augmented generation and targeted retraining.

This creates a virtuous cycle where every uncertain case contributes to future accuracy improvements, progressively reducing manual intervention and expanding the scope of trusted predictions.

## Frequently Asked Questions

### How is the confidence score computed?

The confidence score in Automation is a consensus-based reliability measure, not a simple probability. It is computed by analyzing the level of agreement between multiple models, each trained independently or with complementary strategies, on the same document field.

When these models produce matching or highly similar predictions, the confidence is high. When they disagree significantly, the confidence drops. On top of that, a dedicated arbitration and correction model acts as a referee: it takes all predictions, compares their structural and semantic coherence, and assigns a final confidence level (`low`, `medium`, `high`, or soon `certain`).

### Does Automation introduce additional latency?

Yes, Automation introduces some additional latency, but in most cases, it remains minimal. This is because the ensemble of models used for prediction is executed in parallel, which allows us to keep response times close to those of a single-model pipeline.

However, depending on the number and complexity of models involved, or the document type, the latency can occasionally be a few times longer than a standard call. The tradeoff is intentional: slightly longer processing time in exchange for higher accuracy and richer metadata, including the confidence score.

### What should I do with low confidence extractions?

We recommend routing `Medium` and lower confidence extractions to a human validation layer, or using fallback logic (e.g., default values, user input).

Lower confidence extractions are ideal candidates for feedback-driven improvement via our continuous learning loop using the RAG feature.

### Does Automation work with any type of documents or fields?

Automation is fully compatible with all document types and extracted fields supported by Mindee.\
Every extracted field, whether it's a piece of text, a number, a date, an amount, or any other data type, benefits from the same ensemble evaluation and confidence scoring logic. This consistent approach ensures a uniform and predictable developer experience, regardless of the document format or use case.

Moreover, nested objects and arrays of objects (e.g., `line_items` in invoices or tables in receipts) also receive individual confidence scores per field, enabling fine-grained control over complex data structures.


# Polygons (Bounding Boxes)

Add the polygon coordinates of each extracted field to the API response.

<figure><img src="/files/8HRkzqSIlZSiUJdmZgOd" alt="bounding-box polygons displayed in the Live Interface" width="375"><figcaption></figcaption></figure>

## Overview

The **polygons** option, also commonly referred to as **bounding boxes**, is a feature you can enable in your Models.\
It indicates the precise polygonal area on the document where the value for each extracted field was detected.

* These polygons define the **exact location on the document** containing the extracted data.
* They enable visual verification by highlighting where the model found each value on the document.
* This feature is especially useful for building user interfaces that overlay extracted data on document images for validation or review.

## How It Works <a href="#how-it-works" id="how-it-works"></a>

* A polygon (bounding box) is an array of points outlining a closed shape around the detected data area.
* Points are given as normalized coordinates relative to the document page dimensions:
  * Each coordinate is a float between `0` and `1`.
  * `(0, 0)` corresponds to the top-left corner; `(1, 1)` corresponds to the bottom-right corner.
* Multiple polygons can be returned if there are several extracted fields or data areas.

## **Activate Polygons**

### Activate Polygons on the Platform

When setting the activation state of a feature on the Platform, this will be the default.\
All API calls will use the default state unless explicitly set otherwise during the API call.

This is useful for project managers, as it allows activating or deactivating a feature across all API calls on that model.

Anyone with write access to the model can set the option's default value.

{% @supademo/embed url="<https://app.supademo.com/demo/cmeidsob99f8xh3pybta14v42>" demoId="cmeidsob99f8xh3pybta14v42" %}

### Activate Polygons via API Calls

If you need finer-grained control over when the feature is used, you can activate or deactivate it when making API calls.

This allows dynamically setting the activation state using your internal business or domain logic.

Check the [Client Configuration](/integrations/client-libraries-sdk/configure-the-client#optional-features-configuration) section if using our [Client Libraries / SDKs](/integrations/client-libraries-sdk).

## Use the Polygon Result <a href="#example-polygon-bounding-box-data" id="example-polygon-bounding-box-data"></a>

We highly recommend using our [Client Libraries / SDKs](/integrations/client-libraries-sdk), as they include various geometry functions for ease of processing.

Specifically for handling polygons, take a look at the [Extraction Result](/extraction-models/sdk-integration/extraction-result#locations) section.

Otherwise, take a look at the [Manual Integration](/integrations/api-reference#get-v2-inferences-inference_id) section.

## Some Use Cases <a href="#use-cases" id="use-cases"></a>

* Visual validation by overlaying bounding boxes on document previews for users to confirm extraction accuracy.
* Debugging extraction quality by analyzing detected regions.
* Enhancing UI with highlighted fields for better user experience.


# Raw Text (Full OCR)

Add the full text content of your documents to the API response.

## Overview

By activating this feature, the full text content of the document will be extracted using OCR (Optical Character Recognition) technology.

The resulting content will be included in the API response.

In the response, each page will have its text content filled.

## Use Cases

Having access to the raw text can be particularly useful when:

* needing to save a copy of the text for your records
* perform searches within the text

## Activate Raw Text

### Activate Raw Text via API Calls

If you need finer-grained control over when the feature is used, you can activate or deactivate it when making API calls.

This allows dynamically setting the activation state using your internal business or domain logic.

Check the [Client Configuration](/integrations/client-libraries-sdk/configure-the-client#optional-features-configuration) section if using our [Client Libraries / SDKs](/integrations/client-libraries-sdk).

## Use the Raw Text Result

We highly recommend using our [Client Libraries / SDKs](/integrations/client-libraries-sdk), as they include various functions for ease of processing.

Specifically for handling the raw text, take a look at the [Response Processing](/integrations/client-libraries-sdk/process-the-response#raw-text) section.

Otherwise, take a look at the [Manual Integration](/integrations/api-reference#get-v2-inferences-inference_id) section.


# SDK Integration

Integrate an Extraction model using the Mindee SDKs.

Use the SDKs to send documents to an Extraction model and process dynamic results.

This section helps you choose the right starting point and move through the full integration flow.

### Choose your path

Start with the page that matches your next step:

* [Extraction Quick Start](/extraction-models/sdk-integration/quick-start) ⇒ install a client library, send a file, and get your first result.
* [Extraction Configuration](/extraction-models/sdk-integration/extraction-configuration) ⇒ set model-specific parameters and optional features.
* [Extraction Result Fields](/extraction-models/sdk-integration/extraction-result) ⇒ access dynamic fields and metadata in the SDK response.

### Typical SDK workflow

{% stepper %}
{% step %}

#### Create Your Model

[Create your Extraction Model](/extraction-models/extraction-models-overview#creating-from-scratch) on the Mindee platform.

Upload some samples to the [Live Test](/models/live-test) to validate the model.
{% endstep %}

{% step %}

#### Install and authenticate

Install the client library for your language.

Prepare your API key and initialize the client.
{% endstep %}

{% step %}

#### Send a document

Choose your model ID and send a file or URL for processing.

Start with polling unless you already use webhooks.
{% endstep %}

{% step %}

#### Process the results

Read extracted values from the response fields.

Handle field types based on your Data Schema.
{% endstep %}
{% endstepper %}

### Before you start

Have these ready:

* an API key
* your Extraction Model's unique ID
* a sample document for testing
* a language choice for the SDK

### What is specific to Extraction models

Extraction responses are dynamic.

Field names come from your model's Data Schema. Field types can be simple values, objects, or lists.

Optional metadata such as confidence scores and locations depends on which features are enabled for the request.

For better results, make sure your schema design is solid before tuning request-time options.

### Shared SDK building blocks

Integration builds on the same client library concepts used across all model types.

* [Client Libraries / SDKs](/integrations/client-libraries-sdk)
* [Client Configuration](/integrations/client-libraries-sdk/configure-the-client)
* [Send a File or URL](/integrations/client-libraries-sdk/send-a-file-or-url)
* [Response Processing](/integrations/client-libraries-sdk/process-the-response)
* [Webhook Results](/integrations/webhooks)
* [Error Handling](/integrations/problem-database)


# Extraction Quick Start

Quickest way to get started using the client libraries.

## Installation Instructions

{% tabs %}
{% tab title="Python" %}
Requires Python ≥ 3.9. Python ≥ 3.11 is recommended.

Simply install the [PyPi package](https://pypi.org/project/mindee/) using `pip`:

```sh
pip install -U mindee~=5.2
```

{% endtab %}

{% tab title="Node.js" %}
Requires Node.js ≥ 20.1. Node.js ≥ 22 is recommended.

Simply install the [NPM package](https://www.npmjs.com/package/mindee):

```sh
npm install mindee@^5.5.0
```

{% endtab %}

{% tab title="PHP" %}
Requires PHP ≥ 8.1. PHP ≥ 8.3 is recommended.

Simply install the [Packagist package](https://packagist.org/packages/mindee/mindee) using [composer](https://getcomposer.org/):

```sh
php composer.phar require "mindee/mindee:>=3.0"
```

{% endtab %}

{% tab title="Ruby" %}
Requires Ruby ≥ 3.2.

Simply install the [gem](https://rubygems.org/gems/mindee) using:

```shell
gem install mindee -v '~> 5.2'
```

{% endtab %}

{% tab title="Java" %}
Requires Java ≥ 11. Java ≥ 17 is recommended.

Group ID: `com.mindee.sdk`\
Artifact ID: `mindee-api-java`\
Version: `5.2.0` or greater

There are various installation methods, Maven, Gradle, etc:

[Installation Details](https://central.sonatype.com/artifact/com.mindee.sdk/mindee-api-java)
{% endtab %}

{% tab title=".NET" %}
.NET ≥ 8.0 is recommended.

Simply install the [NuGet package](https://www.nuget.org/packages/Mindee) using `dotnet add`:

```sh
dotnet add package Mindee --version 4.4
```

{% endtab %}
{% endtabs %}

Don't see support for your favorite language or framework? [Make a feature request!](https://feedback.mindee.com/?b=682f69c9e2404756e7e68d1c)

## Send a File and Poll

Make a note of your model's ID for use in the API.

When getting started, we recommend using the polling method which will be quickest (unless you happen to already have access to a public-facing Web server).

Here are basic code examples, these are self-contained and can be run as-is:

{% tabs %}
{% tab title="Python" %}
Requires Python ≥ 3.10. Python ≥ 3.12 is recommended.\
Requires the [Mindee Python client library](https://pypi.org/project/mindee/) version **5.1.1** or greater.

{% code lineNumbers="true" %}

```python
from mindee import PathInput
from mindee.v2 import (
    Client,
    ExtractionParameters,
    ExtractionResponse,
)

input_path = "/path/to/the/file.ext"
api_key = "MY_API_KEY"
model_id = "MY_MODEL_ID"

# Init a new client
mindee_client = Client(api_key)

# Set Extraction parameters
model_params = ExtractionParameters(
    # ID of the model, required.
    model_id=model_id,

    # Options: set to `True` or `False` to override defaults

    # Enhance extraction accuracy with Retrieval-Augmented Generation.
    rag=None,
    # Extract the full text content from the document as strings.
    raw_text=None,
    # Calculate bounding box polygons for all fields.
    polygon=None,
    # Boost the precision and accuracy of all extractions.
    # Calculate confidence scores for all fields.
    confidence=None,
)

# Load a file from disk
input_source = PathInput(input_path)

# Send for processing
response = mindee_client.enqueue_and_get_result(
    ExtractionResponse,
    input_source,
    model_params,
)

# Print a brief summary of the parsed data
print(response.inference)

# Access the result fields
fields: dict = response.inference.result.fields
```

{% endcode %}

Also take a look at the [Extraction Result](https://docs.mindee.com/extraction-models/sdk-integration/extraction-result) documentation.
{% endtab %}

{% tab title="Node.js" %}
Requires Node.js ≥ 20.1. Node.js ≥ 22 is recommended.\
Requires the [Mindee Node.js client library](https://www.npmjs.com/package/mindee/) version **5.5.0** or greater.

{% code lineNumbers="true" %}

```javascript
import * as mindee from "mindee";
// If you're on CommonJS:
// const mindee = require("mindee");

const apiKey = "MY_API_KEY";
const filePath = "/path/to/the/file.ext";
const modelId = "MY_MODEL_ID";

// Init a new client
const mindeeClient = new mindee.Client(
  { apiKey: apiKey }
);

// Set Extraction parameters
const modelParams = {
  modelId: modelId,

  // Options: set to `true` or `false` to override defaults

  // Enhance extraction accuracy with Retrieval-Augmented Generation.
  rag: undefined,
  // Extract the full text content from the document as strings.
  rawText: undefined,
  // Calculate bounding box polygons for all fields.
  polygon: undefined,
  // Boost the precision and accuracy of all extractions.
  // Calculate confidence scores for all fields.
  confidence: undefined,
};

// Load a file from disk
const inputSource = new mindee.PathInput({ inputPath: filePath });

// Send for processing
const response = await mindeeClient.enqueueAndGetResult(
  mindee.product.Extraction,
  inputSource,
  modelParams,
);

// print a string summary
console.log(response.inference.toString());

// Access the result fields
const fields = response.inference.result.fields;
```

{% endcode %}

Also take a look at the [Processing Results](https://docs.mindee.com/extraction-models/sdk-integration/extraction-result) documentation.
{% endtab %}

{% tab title="PHP" %}
Requires PHP ≥ 8.1. PHP ≥ 8.3 is recommended.\
Requires the [Mindee PHP client library](https://packagist.org/packages/mindee/mindee) version **3.0.0** or greater.

{% code lineNumbers="true" %}

```php
<?php

use Mindee\Input\PathInput;
use Mindee\V2\Client;
use Mindee\V2\Product\Extraction\Params\ExtractionParameters;
use Mindee\V2\Product\Extraction\ExtractionResponse;

$apiKey = "MY_API_KEY";
$modelId = "MY_MODEL_ID";
$filePath = "/path/to/the/file.ext";

// Init a new client
$mindeeClient = new Client($apiKey);

// Set Extraction parameters
$modelParams = new ExtractionParameters(
    // ID of the model, required.
    $modelId,

    // Options: set to `true` or `false` to override defaults

    // Enhance extraction accuracy with Retrieval-Augmented Generation.
    rag: null,
    // Extract the full text content from the document as strings.
    rawText: null,
    // Calculate bounding box polygons for all fields.
    polygon: null,
    // Boost the precision and accuracy of all extractions.
    // Calculate confidence scores for all fields.
    confidence: null
);

// Load a file from disk
$inputSource = new PathInput($filePath);

// Send for processing using polling
$response = $mindeeClient->enqueueAndGetResult(
    ExtractionResponse::class,
    $inputSource,
    $modelParams
);

// Print a summary of the response
echo strval($response->inference);

// Access the extracted fields
$fields = $response->inference->result->fields;
```

{% endcode %}

Also take a look at the [Processing Results](https://docs.mindee.com/extraction-models/sdk-integration/extraction-result) documentation.
{% endtab %}

{% tab title="Ruby" %}
Requires Ruby ≥ 3.2.\
Requires the [Mindee Ruby client library](https://rubygems.org/gems/mindee) version **5.2.1** or greater.

{% code lineNumbers="true" %}

```ruby
require 'mindee'
require 'mindee/v2/product'

input_path = '/path/to/the/file.ext'
api_key = 'MY_API_KEY'
model_id = 'MY_MODEL_ID'

# Init a new client
mindee_client = Mindee::V2::Client.new(api_key: api_key)

# Set Extraction parameters
model_params = {
    # ID of the model, required.
    model_id: model_id,

    # Options: set to `true` or `false` to override defaults

    # Enhance extraction accuracy with Retrieval-Augmented Generation.
    rag: nil,
    # Extract the full text content from the document as strings.
    raw_text: nil,
    # Calculate bounding box polygons for all fields.
    polygon: nil,
    # Boost the precision and accuracy of all extractions.
    # Calculate confidence scores for all fields.
    confidence: nil
}

# Load a file from disk
input_source = Mindee::Input::Source::PathInputSource.new(input_path)

# Send for processing
response = mindee_client.enqueue_and_get_result(
    Mindee::V2::Product::Extraction::Extraction,
    input_source,
    model_params
)

# Print a brief summary of the parsed data
puts response.inference

# Access the result fields
fields = response.inference.result.fields

# fields.get_simple_field('my_simple_field')
# fields.get_list_field('my_list_field')
# fields.get_object_field('my_object_field')
```

{% endcode %}

Also take a look at the [Processing Results](https://docs.mindee.com/extraction-models/sdk-integration/extraction-result) documentation.
{% endtab %}

{% tab title="Java" %}
Requires Java ≥ 11. Java ≥ 21 is recommended.\
Requires the [Mindee Java SDK](https://central.sonatype.com/artifact/com.mindee.sdk/mindee-api-java) version **5.2.0** or greater.

{% code lineNumbers="true" %}

```java
import com.mindee.input.LocalInputSource;
import com.mindee.v2.MindeeClient;
import com.mindee.v2.product.extraction.params.ExtractionParameters;
import com.mindee.v2.product.extraction.ExtractionResponse;
import java.io.IOException;

public class SimpleMindeeClientV2 {

  public static void main(String[] args)
      throws IOException, InterruptedException
  {
    String apiKey = "MY_API_KEY";
    String modelId = "MY_MODEL_ID";
    String filePath = "/path/to/the/file.ext";

    // Init a new client
    var mindeeClient = new MindeeClient(apiKey);

    // Set Extraction parameters
    var modelParams = ExtractionParameters
        // ID of the model, required.
        .builder(modelId)

        // Options: set to `true` or `false` to override defaults

        // Enhance extraction accuracy with Retrieval-Augmented Generation.
        .rag(null)
        // Extract the full text content from the document as strings.
        .rawText(null)
        // Calculate bounding box polygons for all fields.
        .polygon(null)
        // Boost the precision and accuracy of all extractions.
        // Calculate confidence scores for all fields.
        .confidence(null)

        .build();

    // Load a file from disk
    var inputSource = new LocalInputSource(filePath);

    // Send for processing using polling
    ExtractionResponse response = mindeeClient.enqueueAndGetResult(
        ExtractionResponse.class,
        inputSource,
        modelParams
    );

    // Print a summary of the response
    System.out.println(response.getInference().toString());

    // Access the result fields
    var fields = response.getInference().getResult().getFields();
  }
}
```

{% endcode %}

Also take a look at the [Processing Results](https://docs.mindee.com/extraction-models/sdk-integration/extraction-result) documentation.
{% endtab %}

{% tab title=".NET" %}
.NET ≥ 8.0 is recommended.\
Requires the [Mindee .NET client library](https://www.nuget.org/packages/Mindee) version **4.3.0** or greater.

{% code lineNumbers="true" %}

```csharp
using Mindee;
using Mindee.Input;
using Mindee.V2;
using Mindee.V2.Product.Extraction;
using Mindee.V2.Product.Extraction.Params;

string filePath = "/path/to/the/file.ext";
string apiKey = "MY_API_KEY";
string modelId = "MY_MODEL_ID";

// Construct a new client
Client mindeeClient = new Client(apiKey);

// Set Extraction parameters
var modelParams = new ExtractionParameters(
    modelId: modelId

    // Options: set to `true` or `false` to override defaults

    // Enhance extraction accuracy with Retrieval-Augmented Generation.
    , rag: null
    // Extract the full text content from the document as strings.
    , rawText: null
    // Calculate bounding box polygons for all fields.
    , polygon: null
    // Boost the precision and accuracy of all extractions.
    // Calculate confidence scores for all fields.
    , confidence: null
);

// Load a file from disk
var inputSource = new LocalInputSource(filePath);

// Upload the file
var response = await mindeeClient.EnqueueAndGetResultAsync<ExtractionResponse>(
    inputSource, modelParams);

// Print a summary of the response
System.Console.WriteLine(response.Inference.ToString());

// Access the extracted fields
var fields = response.Inference.Result.Fields;
```

{% endcode %}

Also take a look at the [Processing Results](https://docs.mindee.com/extraction-models/sdk-integration/extraction-result) documentation.
{% endtab %}
{% endtabs %}

### Details on Sending

For details on available options and advanced usage, check the following sections:

* [Client Configuration](/integrations/client-libraries-sdk/configure-the-client)
* [Load and Adjust a File](/integrations/client-libraries-sdk/load-and-adjust-a-file)
* [Load an URL](/integrations/client-libraries-sdk/load-an-url)
* [Send a File or URL](/integrations/client-libraries-sdk/send-a-file-or-url)

## Process Extraction Results

Once you've sent the file and retrieved the response, you can start accessing the results.

The Extraction model's fields will be in the `fields` object in the return (the `response` variable returned from the above step).

Each key in the `fields` object corresponds to the field's `name` in your Data Schema.

You'll want to adapt your processing depending on the [type of field](/extraction-models/data-schema#field-types), for example when looping over lists or accessing sub-fields.

{% tabs %}
{% tab title="Python" %}
Accessing simple values, using the name of the field in the Data Schema.

You can (should!) specify the type of value, the possible types are `str` , `bool` , `float` .\
Note that all types may be `None`.

```python
from mindee.v2 import ExtractionResponse

def handle_response(response: ExtractionResponse):
  fields = response.inference.result.fields

  # texts, dates, classifications ...
  string_field_value: str | None = fields.get_simple_field("string_field").value

  # a JSON float will be a float
  float_field_value: float | None = fields.get_simple_field("float_field").value

  # even if the API always returns an integer, the type will be float
  int_field_value: float | None = fields.get_simple_field("int_field").value

  # booleans
  bool_field_value: bool | None = fields.get_simple_field("bool_field").value
```

Accessing a list of simple values, where `my_list_field` is the name of the field in the Model.

```python
from mindee.v2 import ExtractionResponse

def handle_response(response: ExtractionResponse):
    fields = response.inference.result.fields

    simple_list_field = fields.get_list_field("my_simple_list_field")

    # Loop over the list of Simple fields
    for list_item in simple_list_field.items:
        item_value = list_item.value
```

Accessing an object field and its sub-fields, where `my_object_field` is the name of the field in the Model. In this hypothetical case, the object has a sub-field named `subfield_1` .

```python
from mindee.v2 import ExtractionResponse

def handle_response(response: ExtractionResponse):
    fields = response.inference.result.fields

    object_field = fields.get_object_field("my_object_field")

    sub_fields: dict = object_field.fields

    # grab a single sub-field
    subfield_1 = sub_fields["subfield_1"]

    # loop over sub-fields
    for field_name, sub_field in sub_fields.items():
        sub_field.value
```

Accessing a list of objects, where `my_object_list_field` is the name of the field in the Model.

```python
from mindee.v2 import ExtractionResponse

def handle_response(response: ExtractionResponse):
    fields = response.inference.result.fields

    object_list_field = fields.get_list_field("my_object_list_field")

    # Loop over the list of Object fields
    for object_item in object_list_field.items:
        # grab a single sub-field
        sub_field_value = object_item.fields["sub_field"].value
        
        # loop over sub-fields
        for field_name, sub_field in object_item.fields.items():
            sub_field.value
```

{% endtab %}

{% tab title="Node.js" %}
Accessing simple values, using the name of the field in the Data Schema.

Access fields as `SimpleField` instances when retrieving their value.

```typescript
handleResponse(response) {
  const fields = response.inference.result.fields;
  
  // texts, dates, classifications ...
  const stringFieldValue = fields.getSimpleField("string_field").stringValue;
  
  // no distinction between floats and integers
  const floatFieldValue = fields.getSimpleField("float_field").numberValue;
  const intFieldValue = fields.getSimpleField("int_field").numberValue;
  
  // booleans
  const boolFieldValue = fields.getSimpleField("bool_field").booleanValue;
}
```

Accessing a list of values, where `my_simple_list_field` is the name of the field in the Model.

We need to specify that the field is a `ListField` in order to access its items.

```javascript
handleResponse(response) {
  const fields = response.inference.result.fields;

  const simpleListField = fields.getListField("my_simple_list_field");

  const simpleItems = simpleListField.simpleItems;

  // Loop over the list of Simple fields
  for (const itemField of simpleItems) {
    // Choose the appropriate accessor:
    // stringValue, numberValue, booleanValue
    const fieldValue = itemField.stringValue;
  }
}
```

Accessing an object field and its sub-fields, where `my_object_field` is the name of the field in the Model. In this hypothetical case, the object has a sub-field named `subfield_1` .

```javascript
handleResponse(response) {
  const fields = response.inference.result.fields;

  const objectField = fields.getObjectField("my_object_field");

  const simpleSubFields = objectField.simpleFields;

  // grab a single sub-field
  const subfield1 = subFields.get("subfield_1");

  // loop over simple sub-fields
  subFields.forEach((simpleSubFields, fieldName) => {
    // Choose the appropriate accessor:
    // stringValue, numberValue, booleanValue
    const fieldValue = subField.stringValue;
  });

  // Object fields can also have lists:
  const listSubFields = objectField.listFields;
}
```

Accessing a list of objects, where `my_object_list_field` is the name of the field in the Model.

We need to specify that the field is a `ListField` in order to access its items.

```javascript
handleResponse(response) {
  const fields = response.inference.result.fields;

  const fieldObjectList = fields.getListField("my_object_list_field");

  const objectItems = fieldSimpleList.objectItems;

  // Loop over the list of Object fields
  for (const itemField of objectItems) {
    const simpleSubFields = itemField.simpleFields;

    // grab a single sub-field
    const subField1 = subFields.get("subfield_1");

    // Choose the appropriate accessor:
    // stringValue, numberValue, booleanValue
    const subFieldValue = subField1.stringValue;

    // loop over simple sub-fields
    simpleSubFields.forEach((subField, fieldName) => {
      // Choose the appropriate accessor:
      // stringValue, numberValue, booleanValue
      const fieldValue = subField.stringValue;
    });

    // Object fields can also have lists:
    const listSubFields = itemField.listFields;
  }
}
```

{% endtab %}

{% tab title="PHP" %}
Accessing simple values, using the name of the field in the Data Schema.

Access fields as `SimpleField` instances when retrieving their value.

```php
use Mindee\V2\Product\Extraction\ExtractionResponse;

public function handleResponse(ExtractionResponse $response)
{
    $fields = $response->inference->result->fields;

    // texts, dates, classifications ...
    $stringFieldValue = $fields->getSimpleField('string_field')->value;

    // a JSON float will be a float
    $floatFieldValue = $fields->getSimpleField('float_field')->value;

    // even if the API always returns an integer, the type will be float
    $intFieldValue = $fields->getSimpleField('int_field')->value;

    // booleans
    $boolFieldValue = $fields->getSimpleField('bool_field')->value;
}
```

Accessing a list of values, where `my_simple_list_field` is the name of the field in the Model.

We need to specify that the field is a `ListField` in order to access its items.

```php
use Mindee\V2\Product\Extraction\ExtractionResponse;

public function handleResponse(ExtractionResponse $response)
{
    $fields = $response->inference->result->fields;

    $simpleListField = $fields->getListField('my_simple_list_field');

    // access a value at a given position
    $fieldFirstValue = $simpleListField->items[0]->value;

    // Loop over the list of Simple fields
    foreach ($simpleListField->items as $listItem) {
        $itemValue = $listItem->value;
    }
}
```

Accessing an object field and its sub-fields, where `my_object_field` is the name of the field in the Model. In this hypothetical case, the object has a sub-field named `subfield_1` .

```php
use Mindee\V2\Product\Extraction\ExtractionResponse;

public function handleResponse(ExtractionResponse $response)
{
    $fields = $response->inference->result->fields;

    $objectField = $fields->getObjectField('my_object_field');
    $subFields = $objectField->fields;

    // grab a single sub-field
    $subfield1 = $subFields->getSimpleField('subfield_1');

    // loop over sub-fields
    foreach ($subFields as $fieldName => $subField) {
        $fieldValue = $subField->value;
    }
}
```

Accessing a list of objects, where `my_object_list_field` is the name of the field in the Model.

We need to specify that the field is a `ListField` in order to access its items.

```php
use Mindee\V2\Product\Extraction\ExtractionResponse;

public function handleResponse(ExtractionResponse $response)
{
    $fields = $response->inference->result->fields;
    $fieldObjectList = $fields->getObjectField('my_object_list_field');

    // access an object at a given position
    $objectItem0 = $fieldObjectList->items[0];
    $subField0Value = $objectItem0->fields->get('sub_field')->value;

    // Loop over the list of Object fields
    foreach ($myObjectListField->items as $objectItem) {
        $subFieldValue = $objectItem->fields->get('sub_field')->value;
    }
}
```

{% endtab %}

{% tab title="Ruby" %}
Accessing simple values, using the name of the field in the Data Schema.

Access fields as `SimpleField` instances when retrieving their value.

```ruby
def handle_response(response)
  fields = response.inference.result.fields

  # texts, dates, classifications ...
  string_field_value = fields.get_simple_field('string_field').string_value

  # a JSON float will be a float
  float_field_value = fields.get_simple_field('float_field').float_value

  # even if the API always returns an integer, the type will be float
  int_field_value = fields.get_simple_field('int_field').float_value

  # booleans
  bool_field_value = fields.get_simple_field('bool_field').boolean_value
end
```

Accessing a list of values, where `my_simple_list_field` is the name of the field in the Model.

Access the list as a `ListField` instance, and the items as `SimpleField` instances.

```ruby
def handle_response(response)
  fields = response.inference.result.fields
  
  my_list_field = fields.get_list_field('my_simple_list_field')

  # access a value at a given position
  simple_list_field = my_list_field.simple_items[0].value

  # Loop over the list of Simple fields
  simple_list_field.simple_items.each do |list_item|
    item_value = list_item.value
  end
end
```

Accessing an object field and its sub-fields, where `my_object_field` is the name of the field in the Model. In this hypothetical case, the object has a sub-field named `subfield_1` .

```ruby
def handle_response(response)
  fields = response.inference.result.fields

  object_field = fields.get_object_field('my_object_field')
  sub_fields = object_field.fields
  
  # grab a single sub-field
  sub_field_value = sub_fields.get_simple_field('sub_field').value
  
  # loop over sub-fields
  sub_fields.each_value do |sub_field|
    field_value = sub_field.value
  end
end
```

Accessing a list of objects, where `my_object_list_field` is the name of the field in the Model.

Access the list as a `ListField` instance, and the items as `ObjectField` instances.

```ruby
def handle_response(response)
  fields = response.inference.result.fields
  
  object_list_field = fields.get_list_field('my_object_list_field')

  # access an object at a given position
  object_item_0 = object_list_field.items[0]
  sub_field_0_value = object_item_0.fields.get('sub_field').value

  # Loop over the list of Object fields
  object_list_field.fields.each do |object_item|
    sub_field_value = object_item.fields.get('sub_field').value
  end
end
```

{% hint style="info" %}
You can technically access all field types by their index: `fields['field_name']`

This is heavily discouraged and **unsupported**.
{% endhint %}
{% endtab %}

{% tab title="Java" %}
Accessing simple values, using the name of the field in the Data Schema.

Access fields as `SimpleField` instances when retrieving their value.

We also need to specify the type of value, the possible types are `String` , `Boolean` , `Double` .\
Note that all types may be `null`.

```java
import com.mindee.v2.product.extraction.ExtractionResponse;

public void handleResponse(ExtractionResponseresponse) {
  var fields = response.getInference().getResult().getFields();

  // texts, dates, classifications ...
  String stringFieldValue = fields.getSimpleField("string_field")
        .getStringValue();

  // a JSON float will be a Double
  Double floatFieldValue = fields.getSimpleField("float_field")
      .getDoubleValue();

  // even if the API always returns an integer, the type will be Double
  Double intFieldValue = fields.getSimpleField("int_field")
      .getDoubleValue();

  // booleans
  Boolean boolFieldValue = fields.getSimpleField("bool_field")
      .getBooleanValue();
}
```

Accessing a list of simple values, where `my_simple_list_field` is the name of the field in the Model.

We need to specify that the field is a `ListField` in order to access its `SimpleItems`.

For each item in the list, we also need to specify the correct field and value type, as described above.

```java
import com.mindee.v2.product.extraction.ExtractionResponse;
import com.mindee.parsing.v2.field.ListField;
import com.mindee.parsing.v2.field.SimpleField;

public void handleResponse(ExtractionResponseresponse) {
  var fields = response.getInference().getResult().getFields();

  ListField simpleListField = fields.getListField("my_simple_list_field");

  List<SimpleField> simpleItems = simpleListField.getSimpleItems();

  // Loop over the list of Simple fields
  for (SimpleField itemField : simpleItems) {
    // Choose the appropriate value type accessor method:
    // String, Double, Boolean
    String fieldValue = itemField.getStringValue();
  }
}
```

Accessing an object field and its sub-fields, where `my_object_field` is the name of the field in the Model. In this hypothetical case, the object has a sub-field named `subfield_1` .

```java
import com.mindee.v2.product.extraction.ExtractionResponse;
import com.mindee.parsing.v2.field.ListField;
import com.mindee.parsing.v2.field.ObjectField;
import com.mindee.parsing.v2.field.SimpleField;

public void handleResponse(ExtractionResponseresponse) {
  var fields = response.getInference().getResult().getFields();

  ObjectField objectField = fields.getObjectField("my_object_field");

  HashMap<String, SimpleField> subFields = objectField.getSimpleFields();

  // grab a single sub-field
  SimpleField subfield1 = subFields.get("subfield_1");

  // loop over sub-fields
  for (Map.Entry<String, SimpleField> entry : subFields.entrySet()) {
    String fieldName = entry.getKey();
    SimpleField subField = entry.getValue();
    
    // Choose the appropriate value type accessor method:
    // String, Double, Boolean
    String fieldValue = subField.getStringValue();
  }
}
```

Accessing a list of objects, where `my_object_list_field` is the name of the field in the Model.

```java
import com.mindee.v2.product.extraction.ExtractionResponse;
import com.mindee.parsing.v2.field.ListField;
import com.mindee.parsing.v2.field.ObjectField;
import com.mindee.parsing.v2.field.SimpleField;

public void handleResponse(ExtractionResponseresponse) {
  var fields = response.getInference().getResult().getFields();

  ListField fieldObjectList = fields.getListField("my_object_list_field");

  List<ObjectField> objectItems = listField.getObjectItems();

  // Loop over the list of Object fields
  for (ObjectField itemField : objectItems) {

    // Object sub-fields will always be Simple fields
    HashMap<String, SimpleField> subFields = itemField.getSimpleFields();

    // grab a single sub-field
    SimpleField subfield1 = subFields.get("subfield_1");
    
    // Choose the appropriate value type accessor method:
    // String, Double, Boolean
    String subFieldValue = subfield1.getStringValue();

    // loop over sub-fields
    for (Map.Entry<String, SimpleField> entry : subFields.entrySet()) {
      String fieldName = entry.getKey();
      SimpleField subField = entry.getValue();
    }
  }
}
```

Depending on your requirements, this can be simplified using various custom methods.
{% endtab %}

{% tab title=".NET" %}
Accessing simple values, using the name of the field in the Data Schema.

Access fields as `SimpleField` instances when retrieving their value.

We also need to specify the type of value, the possible types are `string` , `Boolean` , `Double` .\
Note that all types may be `null`.

```csharp
using Mindee.V2.Parsing.Inference.Field;

public void HandleResponse(ExtractionResponse response)
{
    // texts, dates, classifications ...
    string stringFieldValue = fields["string_field"].SimpleField.Value;

    // a JSON float will be a Double
    Double floatFieldValue = fields["float_field"].SimpleField.Value;

    // even if the API always returns an integer, the type will be Double
    Double intFieldValue = fields["int_field"].SimpleField.Value;

    // booleans
    Boolean boolFieldValue = fields["bool_field"].SimpleField.Value;
}
```

Accessing a list of simple values, where `my_list_field` is the name of the field in the Model.

We need to specify that the field is a `ListField` in order to access its `SimpleItems`.

```csharp
using Mindee.V2.Parsing.Inference.Field;

public void HandleResponse(ExtractionResponse response)
{
    InferenceFields fields = response.Inference.Result.Fields;

    ListField simpleListField = fields["my_simple_list_field"].ListField;

    List<SimpleField> simpleItems = simpleListField.SimpleItems;

    // Loop over the list of Simple fields
    foreach (SimpleField itemField in simpleItems)
    {
        // Choose the appropriate value type:
        // string, Double, Boolean
        string fieldValue = itemField.Value;
    }
}
```

Accessing an object field and its sub-fields, where `my_object_field` is the name of the field in the Model. In this hypothetical case, the object has a sub-field named `subfield_1` .

```csharp
using Mindee.V2.Parsing.Inference.Field;

public void HandleResponse(ExtractionResponse response)
{
    InferenceFields fields = response.Inference.Result.Fields;

    ObjectField objectField = fields["my_object_field"].ObjectField;

    Dictionary<string, SimpleField> subFields = objectField.SimpleFields;

    // grab a single sub-field
    SimpleField subField1 = subFields["subfield_1"];

    // loop over sub-fields
    // Note: in C#14, 'field' is a reserved keyword, hence use of 'entry'.
    foreach (KeyValuePair<string, SimpleField> entry in subFields)
    {
        string fieldName = entry.Key;
        SimpleField subField = entry.Value;
        
        // process the SimpleField ...
    }
}
```

Accessing a list of objects, where `my_object_list_field` is the name of the field in the Model.

```csharp
using Mindee.V2.Parsing.Inference.Field;

public void HandleResponse(ExtractionResponse response)
{
    InferenceFields fields = response.Inference.Result.Fields;

    ListField fieldObjectList = fields["my_simple_list_field"].ListField;

    List<ObjectField> objectItems = fieldObjectList.ObjectItems;

    // Loop over the list of Object fields
    foreach (ObjectField itemField in objectItems)
    {
        Dictionary<string, SimpleField> subFields = objectField.SimpleFields;
    
        // grab a single sub-field
        SimpleField subField1 = subFields["subfield_1"];
        
        // Choose the appropriate value type:
        // string, Double, Boolean
        string subFieldValue = subField1.Value;
    
        // loop over sub-fields
        // Note: in C#14, 'field' is a reserved keyword, hence use of 'entry'.
        foreach (KeyValuePair<string, SimpleField> entry in subFields)
        {
            string fieldName = entry.Key;
            SimpleField subField = entry.Value;
        }
    }
}
```

{% endtab %}
{% endtabs %}

### Details on Response Processing

For more details on using the result fields in your application: [Extraction Result](/extraction-models/sdk-integration/extraction-result)

For details on response metadata: [Response Processing](/integrations/client-libraries-sdk/process-the-response)


# Extraction Configuration

Configuration parameters specific to Extraction models.

There are also [Basic Model Configuration](/integrations/client-libraries-sdk/basic-model-configuration) which can be used with all models.

## Optional Features Configuration

Enable or disable [Optional Features](/extraction-models/optional-features).

{% hint style="warning" icon="money-check-dollar-pen" %}
Enabling a feature not in your plan will result in a Payment Required error (HTTP 402).

Check the [Plans and Credits](/account-management/plans#feature-comparison) section for more information.
{% endhint %}

The default activation states for Optional Features are set on the platform.\
Any values set here will override the defaults.

Leave empty or null to use the default platform values.

For example: if the Polygon feature is enabled on the platform, and polygon is explicitly set to `false` in the parameters ⇒ the Polygon feature will **not** be enabled for the API call.

{% tabs %}
{% tab title="Python" %}
Only the `model_id` is required.

```python
model_params = ExtractionParameters(
    # ID of the model, required.
    model_id="MY_MODEL_ID",

    # Optional Features: set to `True` or `False` to override defaults

    # Enhance extraction accuracy with Retrieval-Augmented Generation.
    rag=None,
    # Extract the full text content from the document as strings.
    raw_text=None,
    # Calculate bounding box polygons for all fields.
    polygon=None,
    # Boost the precision and accuracy of all extractions.
    # Calculate confidence scores for all fields.
    confidence=None,
    
    # ... any other options ...
)
```

{% endtab %}

{% tab title="Node.js" %}
Only the `modelId` is required.

```typescript
const modelParams = {
  // ID of the model, required.
  modelId: "MY_MODEL_ID",

  // Optional Features: set to `true` or `false` to override defaults

  // Enhance extraction accuracy with Retrieval-Augmented Generation.
  rag: undefined,
  // Extract the full text content from the document as strings.
  rawText: undefined,
  // Calculate bounding box polygons for all fields.
  polygon: undefined,
  // Boost the precision and accuracy of all extractions.
  // Calculate confidence scores for all fields.
  confidence: undefined,
  
  // ... any other options ...
};
```

{% endtab %}

{% tab title="PHP" %}
Only the `modelId` is required.

```php
$modelParams = new ExtractionParameters(
    // ID of the model, required.
    "MY_MODEL_ID",

    // Optional Features: set to `true` or `false` to override defaults

    // Enhance extraction accuracy with Retrieval-Augmented Generation.
    rag: null,
    // Extract the full text content from the document as strings.
    rawText: null,
    // Calculate bounding box polygons for all fields.
    polygon: null,
    // Boost the precision and accuracy of all extractions.
    // Calculate confidence scores for all fields.
    confidence: null,
    
    // ... any other options ...
);
```

{% endtab %}

{% tab title="Ruby" %}
Only the `model_id` is required.

```ruby
model_params = {
    # ID of the model, required.
    model_id: 'MY_MODEL_ID',

    # Options: set to `true` or `false` to override defaults

    # Enhance extraction accuracy with Retrieval-Augmented Generation.
    rag: nil,
    # Extract the full text content from the document as strings.
    raw_text: nil,
    # Calculate bounding box polygons for all fields.
    polygon: nil,
    # Boost the precision and accuracy of all extractions.
    # Calculate confidence scores for all fields.
    confidence: nil,
    
    # ... any other options ...
}
```

{% endtab %}

{% tab title="Java" %}
Only the `modelId` is required.

```java
var modelParams = ExtractionParameters
    // ID of the model, required.
    .builder("MY_MODEL_ID")

    // Optional Features: set to `true` or `false` to override defaults

    // Enhance extraction accuracy with Retrieval-Augmented Generation.
    .rag(null)
    // Extract the full text content from the document as strings.
    .rawText(null)
    // Calculate bounding box polygons for all fields.
    .polygon(null)
    // Boost the precision and accuracy of all extractions.
    // Calculate confidence scores for all fields.
    .confidence(null)
    
    // ... any other options ...

    // complete the builder
    .build();
```

{% endtab %}

{% tab title=".NET" %}
Only the `modelId` is required.

```csharp
var modelParams = new ExtractionParameters(
    // ID of the model, required.
    modelId: "MY_MODEL_ID"

    // Optional Features: set to `true` or `false` to override defaults

    // Enhance extraction accuracy with Retrieval-Augmented Generation.
    , rag: null
    // Extract the full text content from the document as strings.
    , rawText: null
    // Calculate bounding box polygons for all fields.
    , polygon: null
    // Boost the precision and accuracy of all extractions.
    // Calculate confidence scores for all fields.
    , confidence: null
    
    // ... any other options ...
);
```

{% endtab %}
{% endtabs %}

## Dynamic Model Options

These options allow changing how the model performs an inference on a **per-call basis**.

These features can **only** be used via API.

{% hint style="info" %}
These advanced features are not meant for improving the  model's **overall** accuracy.

Instead, make sure the Data Schema has been [properly optimized](/extraction-models/data-schema#performance-optimization).
{% endhint %}

### Text Context

Give additional guidelines to the model to help it better process a specific document.

Useful when you have important context on the document, **and** when there isn't sufficient information on the document itself to provide that context to the model.

This is a free-form text format.

As an example, you could remove ambiguity for country or regional differences:

"The parts supplier is in Canada, these amounts are in CAD", if there is no address on the document.

### Data Schema

Allows changing the Data Schema on a per-call basis: directly modify the Data Schema: add, remove, or change fields.

The typical use case is when the data needing to be extracted change based on internal business logic.

To download the JSON string appropriate for your model:

1. Go to your model's page
2. On the left-hand menu, click on "General Settings"
3. Scroll down to the "Actions" section
4. Click on the "Download Data Schema" button:<br>

   <figure><img src="/files/OXM8QOI0EXGc5KuXi0MI" alt="The &#x22;Download Data Schema&#x22; button" width="530"><figcaption></figcaption></figure>

### Code Sample

The Data Schema can be passed as a JSON string or by instantiating the appropriate classes.

If passed as a JSON string, it will be validated in the client before being sent to the server.

{% tabs %}
{% tab title="Python" %}
Only the `model_id` is required.

```python
model_params = ExtractionParameters(
    # ID of the model, required.
    model_id="MY_MODEL_ID",

    # Text Context
    text_context="this is an invoice.",

    # Data Schema
    data_schema="{ ... JSON DATA ... }",

    # ... any other options ...
)
```

{% endtab %}

{% tab title="Node.js" %}
Only the `modelId` is required.

```typescript
const modelParams = {
  // ID of the model, required.
  modelId: "MY_MODEL_ID",

  // Text Context
  textContext: "this is an invoice.",

  // Data Schema
  dataSchema: "{ ... JSON DATA ... }",

  // ... any other options ...
};
```

{% endtab %}

{% tab title="PHP" %}
Only the `modelId` is required.

```php
$modelParams = new ExtractionParameters(
    // ID of the model, required.
    "MY_MODEL_ID",

    // Text Context
    textContext: "this is an invoice.",

    // Data Schema
    dataSchema: "{ ... JSON DATA ... }",

    // ... any other options ...
);
```

{% endtab %}

{% tab title="Ruby" %}
Only the `model_id` is required.

```ruby
model_params = {
    # ID of the model, required.
    model_id: 'MY_MODEL_ID',

    # Text Context
    text_context: "this is an invoice.",

    # Data Schema
    data_schema: "{ ... JSON DATA ... }",

    # ... any other options ...
}
```

{% endtab %}

{% tab title="Java" %}
Only the `modelId` is required.

```java
var modelParams = ExtractionParameters
    // ID of the model, required.
    .builder("MY_MODEL_ID")

    // Text Context
    .textContext("this is an invoice.")

    // Data Schema
    .dataSchema("{ ... JSON DATA ... }")

    // ... any other options ...

    // complete the builder
    .build();
```

{% endtab %}

{% tab title=".NET" %}
Only the `modelId` is required.

```csharp
var modelParams = new ExtractionParameters(
    // ID of the model, required.
    modelId: "MY_MODEL_ID"
    
    // Text Context
    , textContext: "this is an invoice."
    
    // Data Schema
    , dataSchema: "{ ... JSON DATA ... }"
    
    // ... any other options ...
);
```

{% endtab %}
{% endtabs %}


# Extraction Result

Reference documentation on processing an Extraction result using the Mindee SDKs.

You'll need to have a response object as described in the [Response Processing](/integrations/client-libraries-sdk/process-the-response) section.

## Accessing Data Schema Fields

Fields are completely dynamic and depend on your model's [Data Schema Overview](/extraction-models/data-schema).

In the client library, you'll have access to the various fields as a key-value mapping type (Python's `dict`, Java's `HashMap`, etc).

Accessing a field is done via its name in the Data Schema.

Each field will be one of the following types:

* A single value, `SimpleField` class.
* A nested object (sub-fields), `ObjectField` class.
* A list or array of fields, `ListField` class.

## `SimpleField` - Single-Value Field

Basic field type having the `value` attribute.\
See the [#value](#value "mention") section below.

In addition, the `Simplefield` class has [#confidence](#confidence "mention") and [#locations](#locations "mention") attributes.

{% tabs %}
{% tab title="Python" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```python
from mindee.v2 import ExtractionResponse

def handle_response(response: ExtractionResponse):
    fields = response.inference.result.fields

    my_simple_field = fields.get_simple_field("my_simple_field")
```

{% endtab %}

{% tab title="Node.js" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```javascript
handleResponse(response) {
  const fields = response.inference.result.fields;

  const simpleField = fields.getSimpleField("my_simple_field");
}
```

{% endtab %}

{% tab title="PHP" %}
Using the `$response` deserialized object from either the polling response or a webhook payload.

```php
use Mindee\V2\Product\Extraction\ExtractionResponse;

public function handleResponse(ExtractionResponse $response)
{
    $fields = $response->inference->result->fields;

    $simpleField = $fields->getSimpleField('my_simple_field');
}
```

{% endtab %}

{% tab title="Ruby" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```ruby
def handle_response(response)
  fields = response.inference.result.fields

  simple_field = fields.get_simple_field('my_simple_field')
end
```

{% endtab %}

{% tab title="Java" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```java
import com.mindee.v2.product.extraction.ExtractionResponse;
import com.mindee.parsing.v2.field.SimpleField;

public void handleResponse(ExtractionResponseresponse) {
  var fields = response.getInference().getResult().getFields();

  SimpleField simpleField = fields.getSimpleField("my_simple_field");
}
```

{% endtab %}

{% tab title=".NET" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```csharp
using Mindee.V2.Product.Extraction;
using Mindee.V2.Parsing.Inference.Field;

public void HandleResponse(ExtractionResponse response)
{
    InferenceFields fields = response.Inference.Result.Fields;

    SimpleField mySimpleField = fields["my_simple_field"].SimpleField;
}
```

{% endtab %}
{% endtabs %}

### `value`

The extracted data value.\
Possible types: string, number (integer or floating-point), boolean.\
All types can be null.

On the platform, you can specify date and classification types.\
These are returned as strings.

For statically-typed languages (C#, Java), the client library will always return a nullable `double` for number values.

{% tabs %}
{% tab title="Python" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```python
from mindee.v2 import ExtractionResponse

def handle_response(response: ExtractionResponse):
  fields = response.inference.result.fields

  # texts, dates, classifications ...
  string_field_value: str | None = fields.get_simple_field("string_field").value

  # a JSON float will be a float
  float_field_value: float | None = fields.get_simple_field("float_field").value

  # even if the API always returns an integer, the type will be float
  int_field_value: float | None = fields.get_simple_field("int_field").value

  # booleans
  bool_field_value: bool | None = fields.get_simple_field("bool_field").value
```

{% endtab %}

{% tab title="Node.js" %}
The `value` attribute is an `Object` type under the hood.

You should use the explicitly-typed accessors, this is recommended for clarity.\
Take a look at your Data Schema to know which typed accessor to use.

```typescript
handleResponse(response) {
  const fields = response.inference.result.fields;
  
  // texts, dates, classifications ...
  const stringFieldValue = fields.getSimpleField("string_field").stringValue;
  
  // no distinction between floats and integers
  const floatFieldValue = fields.getSimpleField("float_field").numberValue;
  const intFieldValue = fields.getSimpleField("int_field").numberValue;
  
  // booleans
  const boolFieldValue = fields.getSimpleField("bool_field").booleanValue;
}
```

If the wrong accessor type is used, an exception will be thrown, something like this:

```
"Value is not a number"
```

{% endtab %}

{% tab title="PHP" %}
Using the `$response` deserialized object from either the polling response or a webhook payload.

```php
use Mindee\V2\Product\Extraction\ExtractionResponse;

public function handleResponse(ExtractionResponse $response)
{
    $fields = $response->inference->result->fields;

    // texts, dates, classifications ...
    $stringFieldValue = $fields->getSimpleField('string_field')->value;

    // a JSON float will be a float
    $floatFieldValue = $fields->getSimpleField('float_field')->value;

    // even if the API always returns an integer, the type will be float
    $intFieldValue = $fields->getSimpleField('int_field')->value;

    // booleans
    $boolFieldValue = $fields->getSimpleField('bool_field')->value;
}
```

{% endtab %}

{% tab title="Ruby" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```ruby
def handle_response(response)
  fields = response.inference.result.fields

  # texts, dates, classifications ...
  string_field_value = fields.get_simple_field('string_field').string_value

  # a JSON float will be a float
  float_field_value = fields.get_simple_field('float_field').float_value

  # even if the API always returns an integer, the type will be float
  int_field_value = fields.get_simple_field('int_field').float_value

  # booleans
  bool_field_value = fields.get_simple_field('bool_field').boolean_value
end
```

{% endtab %}

{% tab title="Java" %}
The `value` attribute is an `Object` type under the hood.

You'll need to explicitly declare the type, otherwise the code will likely not compile.\
Take a look at your Data Schema to know which type to declare.

```java
import com.mindee.v2.product.extraction.ExtractionResponse;

public void handleResponse(ExtractionResponseresponse) {
  var fields = response.getInference().getResult().getFields();

  // texts, dates, classifications ...
  String stringFieldValue = fields.getSimpleField("string_field")
        .getStringValue();

  // a JSON float will be a Double
  Double floatFieldValue = fields.getSimpleField("float_field")
      .getDoubleValue();

  // even if the API always returns an integer, the type will be Double
  Double intFieldValue = fields.getSimpleField("int_field")
      .getDoubleValue();

  // booleans
  Boolean boolFieldValue = fields.getSimpleField("bool_field")
      .getBooleanValue();
}
```

If the wrong type method is used, an exception will be thrown, something like this:

```
ClassCast class java.lang.String cannot be cast to class java.lang.Double
```

{% endtab %}

{% tab title=".NET" %}
The `Value` attribute is a `dynamic` type under the hood.

You should explicitly declare the type, this is recommended for clarity.\
Take a look at your Data Schema to know which type to declare.

```csharp
using Mindee.V2.Parsing.Inference.Field;

public void HandleResponse(ExtractionResponse response)
{
    // texts, dates, classifications ...
    string stringFieldValue = fields["string_field"].SimpleField.Value;

    // a JSON float will be a Double
    Double floatFieldValue = fields["float_field"].SimpleField.Value;

    // even if the API always returns an integer, the type will be Double
    Double intFieldValue = fields["int_field"].SimpleField.Value;

    // booleans
    Boolean boolFieldValue = fields["bool_field"].SimpleField.Value;
}
```

If the wrong type is declared, an exception will be raised, something like this:

```
RuntimeBinderException : Cannot implicitly convert type 'string' to 'double'
```

{% endtab %}
{% endtabs %}

## `ObjectField` - Nested Object Field

Field having a `fields` attribute which is a hash table (Python's `dict`, Java's `HashMap`, etc) of sub-fields.\
See the [#fields](#fields "mention") section below.

In addition, the `ObjectField` class has [#confidence](#confidence "mention") and [#locations](#locations "mention") attributes.

{% tabs %}
{% tab title="Python" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```python
from mindee.v2 import ExtractionResponse

def handle_response(response: ExtractionResponse):
    fields = response.inference.result.fields

    object_field = fields.get_object_field("my_object_field")
```

{% endtab %}

{% tab title="Node.js" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```javascript
handleResponse(response) {
  const fields = response.inference.result.fields;

  const objectField = fields.getObjectField("my_object_field");
}
```

{% endtab %}

{% tab title="PHP" %}
Using the `$response` deserialized object from either the polling response or a webhook payload.

```php
use Mindee\V2\Product\Extraction\ExtractionResponse;

public function handleResponse(ExtractionResponse $response)
{
    $fields = $response->inference->result->fields;

    $objectField = $fields->getObjectField('my_object_field');
}
```

{% endtab %}

{% tab title="Java" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```java
import com.mindee.v2.product.extraction.ExtractionResponse;
import com.mindee.parsing.v2.field.ObjectField;

public void handleResponse(ExtractionResponseresponse) {
  var fields = response.getInference().getResult().getFields();

  ObjectField objectField = fields.getObjectField("my_object_field");
}
```

{% endtab %}

{% tab title=".NET" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```csharp
using Mindee.V2.Parsing.Inference.Field;

public void HandleResponse(ExtractionResponse response)
{
    InferenceFields fields = response.Inference.Result.Fields;

    ObjectField myObjectField = fields["my_object_field"].ObjectField;
}
```

{% endtab %}
{% endtabs %}

### `fields`

The sub-fields as a key-value mapping type (Python `dict`, Java `HashMap`, etc).

Accessing a sub-field is done via its name in the Data Schema.

Each sub-field will be a [#single-value-field-simplefield](#single-value-field-simplefield "mention")or a [#listfield-list-of-fields](#listfield-list-of-fields "mention").

{% tabs %}
{% tab title="Python" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```python
from mindee.v2 import ExtractionResponse

def handle_response(response: ExtractionResponse):
    fields = response.inference.result.fields

    object_field = fields.get_object_field("my_object_field")

    sub_fields: dict = object_field.fields

    # grab a single sub-field
    subfield_1 = sub_fields["subfield_1"]

    # loop over sub-fields
    for field_name, sub_field in sub_fields.items():
        sub_field.value
```

{% endtab %}

{% tab title="Node.js" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```javascript
handleResponse(response) {
  const fields = response.inference.result.fields;

  const objectField = fields.getObjectField("my_object_field");

  const simpleSubFields = objectField.simpleFields;

  // grab a single sub-field
  const subfield1 = subFields.get("subfield_1");

  // loop over simple sub-fields
  subFields.forEach((simpleSubFields, fieldName) => {
    // Choose the appropriate accessor:
    // stringValue, numberValue, booleanValue
    const fieldValue = subField.stringValue;
  });

  // Object fields can also have lists:
  const listSubFields = objectField.listFields;
}
```

{% endtab %}

{% tab title="PHP" %}
Using the `$response` deserialized object from either the polling response or a webhook payload.

```php
use Mindee\V2\Product\Extraction\ExtractionResponse;

public function handleResponse(ExtractionResponse $response)
{
    $fields = $response->inference->result->fields;

    $objectField = $fields->getObjectField('my_object_field');
    $subFields = $objectField->fields;

    // grab a single sub-field
    $subfield1 = $subFields->getSimpleField('subfield_1');

    // loop over sub-fields
    foreach ($subFields as $fieldName => $subField) {
        $fieldValue = $subField->value;
    }
}
```

{% endtab %}

{% tab title="Ruby" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```ruby
def handle_response(response)
  fields = response.inference.result.fields

  object_field = fields.get_object_field('my_object_field')
  sub_fields = object_field.fields
  
  # grab a single sub-field
  sub_field_value = sub_fields.get_simple_field('sub_field').value
  
  # loop over sub-fields
  sub_fields.each_value do |sub_field|
    field_value = sub_field.value
  end
end
```

{% endtab %}

{% tab title="Java" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```java
import com.mindee.v2.product.extraction.ExtractionResponse;
import com.mindee.parsing.v2.field.ListField;
import com.mindee.parsing.v2.field.ObjectField;
import com.mindee.parsing.v2.field.SimpleField;

public void handleResponse(ExtractionResponseresponse) {
  var fields = response.getInference().getResult().getFields();

  ObjectField objectField = fields.getObjectField("my_object_field");

  HashMap<String, SimpleField> subFields = objectField.getSimpleFields();

  // grab a single sub-field
  SimpleField subfield1 = subFields.get("subfield_1");

  // loop over sub-fields
  for (Map.Entry<String, SimpleField> entry : subFields.entrySet()) {
    String fieldName = entry.getKey();
    SimpleField subField = entry.getValue();
    
    // Choose the appropriate value type accessor method:
    // String, Double, Boolean
    String fieldValue = subField.getStringValue();
  }
}
```

{% endtab %}

{% tab title=".NET" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```csharp
using Mindee.V2.Parsing.Inference.Field;

public void HandleResponse(ExtractionResponse response)
{
    InferenceFields fields = response.Inference.Result.Fields;

    ObjectField objectField = fields["my_object_field"].ObjectField;

    Dictionary<string, SimpleField> subFields = objectField.SimpleFields;

    // grab a single sub-field
    SimpleField subField1 = subFields["subfield_1"];

    // loop over sub-fields
    // Note: in C#14, 'field' is a reserved keyword, hence use of 'entry'.
    foreach (KeyValuePair<string, SimpleField> entry in subFields)
    {
        string fieldName = entry.Key;
        SimpleField subField = entry.Value;
        
        // process the SimpleField ...
    }
}
```

{% endtab %}
{% endtabs %}

## `ListField` - List of Fields

Field having an `items` attribute which is a list of fields.\
See the [#items](#items "mention") section below.

In addition, the `ListField` class has a [#confidence](#confidence "mention") attribute.

{% tabs %}
{% tab title="Python" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```python
from mindee.v2 import ExtractionResponse

def handle_response(response: ExtractionResponse):
    fields = response.inference.result.fields

    my_list_field = fields.get_list_field("my_list_field")
```

{% endtab %}

{% tab title="Node.js" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```javascript
handleResponse(response) {
  const fields = response.inference.result.fields;

  const listField = fields.getListField("my_list_field");
}
```

{% endtab %}

{% tab title="PHP" %}
Using the `$response` deserialized object from either the polling response or a webhook payload.

```php
use Mindee\V2\Product\Extraction\ExtractionResponse;

public function handleResponse(ExtractionResponse $response)
{
    $fields = $response->inference->result->fields;
    
    $listField = $fields->getListField('my_list_field');
}
```

{% endtab %}

{% tab title="Java" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```java
import com.mindee.v2.product.extraction.ExtractionResponse;
import com.mindee.parsing.v2.field.ListField;

public void handleResponse(ExtractionResponseresponse) {
  var fields = response.getInference().getResult().getFields();

  ListField listField = fields.getListField("my_list_field");
}
```

{% endtab %}

{% tab title=".NET" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```csharp
using Mindee.V2.Parsing.Inference.Field;

public void HandleResponse(ExtractionResponse response)
{
    InferenceFields fields = response.Inference.Result.Fields;

    ListField myListField = fields["my_list_field"].ListField;
}
```

{% endtab %}
{% endtabs %}

### `items`

List of fields as a variable-length array type (Python `list`, JavaScript `Array`, Java `List`, etc).

Each item in the list will be one of:

* [#simplefield-single-value-field](#simplefield-single-value-field "mention")
* [#objectfield-nested-object-field](#objectfield-nested-object-field "mention")

There will **not** be a mix of both types in the same list.

#### List of `SimpleField`

{% tabs %}
{% tab title="Python" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```python
from mindee.v2 import ExtractionResponse

def handle_response(response: ExtractionResponse):
    fields = response.inference.result.fields

    simple_list_field = fields.get_list_field("my_simple_list_field")

    # Loop over the list of Simple fields
    for list_item in simple_list_field.items:
        item_value = list_item.value
```

{% endtab %}

{% tab title="Node.js" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```javascript
handleResponse(response) {
  const fields = response.inference.result.fields;

  const simpleListField = fields.getListField("my_simple_list_field");

  const simpleItems = simpleListField.simpleItems;

  // Loop over the list of Simple fields
  for (const itemField of simpleItems) {
    // Choose the appropriate accessor:
    // stringValue, numberValue, booleanValue
    const fieldValue = itemField.stringValue;
  }
}
```

{% endtab %}

{% tab title="PHP" %}
Using the `$response` deserialized object from either the polling response or a webhook payload.

```php
use Mindee\V2\Product\Extraction\ExtractionResponse;

public function handleResponse(ExtractionResponse $response)
{
    $fields = $response->inference->result->fields;

    $simpleListField = $fields->getListField('my_simple_list_field');

    // access a value at a given position
    $fieldFirstValue = $simpleListField->items[0]->value;

    // Loop over the list of Simple fields
    foreach ($simpleListField->items as $listItem) {
        $itemValue = $listItem->value;
    }
}
```

{% endtab %}

{% tab title="Ruby" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```ruby
def handle_response(response)
  fields = response.inference.result.fields
  
  my_list_field = fields.get_list_field('my_simple_list_field')

  # access a value at a given position
  simple_list_field = my_list_field.simple_items[0].value

  # Loop over the list of Simple fields
  simple_list_field.simple_items.each do |list_item|
    item_value = list_item.value
  end
end
```

{% endtab %}

{% tab title="Java" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```java
import com.mindee.v2.product.extraction.ExtractionResponse;
import com.mindee.parsing.v2.field.ListField;
import com.mindee.parsing.v2.field.SimpleField;

public void handleResponse(ExtractionResponseresponse) {
  var fields = response.getInference().getResult().getFields();

  ListField simpleListField = fields.getListField("my_simple_list_field");

  List<SimpleField> simpleItems = simpleListField.getSimpleItems();

  // Loop over the list of Simple fields
  for (SimpleField itemField : simpleItems) {
    // Choose the appropriate value type accessor method:
    // String, Double, Boolean
    String fieldValue = itemField.getStringValue();
  }
}
```

{% endtab %}

{% tab title=".NET" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```csharp
using Mindee.V2.Parsing.Inference.Field;

public void HandleResponse(ExtractionResponse response)
{
    InferenceFields fields = response.Inference.Result.Fields;

    ListField simpleListField = fields["my_simple_list_field"].ListField;

    List<SimpleField> simpleItems = simpleListField.SimpleItems;

    // Loop over the list of Simple fields
    foreach (SimpleField itemField in simpleItems)
    {
        // Choose the appropriate value type:
        // string, Double, Boolean
        string fieldValue = itemField.Value;
    }
}
```

{% endtab %}
{% endtabs %}

#### List of `ObjectField`

{% tabs %}
{% tab title="Python" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```python
from mindee.v2 import ExtractionResponse

def handle_response(response: ExtractionResponse):
    fields = response.inference.result.fields

    object_list_field = fields.get_list_field("my_object_list_field")

    # Loop over the list of Object fields
    for object_item in object_list_field.items:
        # grab a single sub-field
        sub_field_value = object_item.fields["sub_field"].value
        
        # loop over sub-fields
        for field_name, sub_field in object_item.fields.items():
            sub_field.value
```

{% endtab %}

{% tab title="Node.js" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```javascript
handleResponse(response) {
  const fields = response.inference.result.fields;

  const fieldObjectList = fields.getListField("my_object_list_field");

  const objectItems = fieldSimpleList.objectItems;

  // Loop over the list of Object fields
  for (const itemField of objectItems) {
    const simpleSubFields = itemField.simpleFields;

    // grab a single sub-field
    const subField1 = subFields.get("subfield_1");

    // Choose the appropriate accessor:
    // stringValue, numberValue, booleanValue
    const subFieldValue = subField1.stringValue;

    // loop over simple sub-fields
    simpleSubFields.forEach((subField, fieldName) => {
      // Choose the appropriate accessor:
      // stringValue, numberValue, booleanValue
      const fieldValue = subField.stringValue;
    });

    // Object fields can also have lists:
    const listSubFields = itemField.listFields;
  }
}
```

{% endtab %}

{% tab title="PHP" %}
Using the `$response` deserialized object from either the polling response or a webhook payload.

```php
use Mindee\V2\Product\Extraction\ExtractionResponse;

public function handleResponse(ExtractionResponse $response)
{
    $fields = $response->inference->result->fields;
    $fieldObjectList = $fields->getObjectField('my_object_list_field');

    // access an object at a given position
    $objectItem0 = $fieldObjectList->items[0];
    $subField0Value = $objectItem0->fields->get('sub_field')->value;

    // Loop over the list of Object fields
    foreach ($myObjectListField->items as $objectItem) {
        $subFieldValue = $objectItem->fields->get('sub_field')->value;
    }
}
```

{% endtab %}

{% tab title="Ruby" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```ruby
def handle_response(response)
  fields = response.inference.result.fields
  
  object_list_field = fields.get_list_field('my_object_list_field')

  # access an object at a given position
  object_item_0 = object_list_field.items[0]
  sub_field_0_value = object_item_0.fields.get('sub_field').value

  # Loop over the list of Object fields
  object_list_field.fields.each do |object_item|
    sub_field_value = object_item.fields.get('sub_field').value
  end
end
```

{% endtab %}

{% tab title="Java" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```java
import com.mindee.v2.product.extraction.ExtractionResponse;
import com.mindee.parsing.v2.field.ListField;
import com.mindee.parsing.v2.field.ObjectField;
import com.mindee.parsing.v2.field.SimpleField;

public void handleResponse(ExtractionResponseresponse) {
  var fields = response.getInference().getResult().getFields();

  ListField fieldObjectList = fields.getListField("my_object_list_field");

  List<ObjectField> objectItems = listField.getObjectItems();

  // Loop over the list of Object fields
  for (ObjectField itemField : objectItems) {

    // Object sub-fields will always be Simple fields
    HashMap<String, SimpleField> subFields = itemField.getSimpleFields();

    // grab a single sub-field
    SimpleField subfield1 = subFields.get("subfield_1");
    
    // Choose the appropriate value type accessor method:
    // String, Double, Boolean
    String subFieldValue = subfield1.getStringValue();

    // loop over sub-fields
    for (Map.Entry<String, SimpleField> entry : subFields.entrySet()) {
      String fieldName = entry.getKey();
      SimpleField subField = entry.getValue();
    }
  }
}
```

{% endtab %}

{% tab title=".NET" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```csharp
using Mindee.V2.Parsing.Inference.Field;

public void HandleResponse(ExtractionResponse response)
{
    InferenceFields fields = response.Inference.Result.Fields;

    ListField fieldObjectList = fields["my_simple_list_field"].ListField;

    List<ObjectField> objectItems = fieldObjectList.ObjectItems;

    // Loop over the list of Object fields
    foreach (ObjectField itemField in objectItems)
    {
        Dictionary<string, SimpleField> subFields = objectField.SimpleFields;
    
        // grab a single sub-field
        SimpleField subField1 = subFields["subfield_1"];
        
        // Choose the appropriate value type:
        // string, Double, Boolean
        string subFieldValue = subField1.Value;
    
        // loop over sub-fields
        // Note: in C#14, 'field' is a reserved keyword, hence use of 'entry'.
        foreach (KeyValuePair<string, SimpleField> entry in subFields)
        {
            string fieldName = entry.Key;
            SimpleField subField = entry.Value;
        }
    }
}
```

{% endtab %}
{% endtabs %}

## Optional Field Attributes

These field attributes are only filled when their respective features are activated.

The attributes are always present even when not activated.

### `confidence`

The confidence level of the extracted value.

The data are only filled if the [Confidence Score and Accuracy Boost](/extraction-models/optional-features/automation-confidence-score) feature is activated.

The instance property is always present, however if the feature is not activated, it will always be empty (the exact type depends on language used: `null`, `undefined`, `None`, etc)

The attribute value will be a one of: `Certain`, `High`, `Medium`, `Low` .\
The language-appropriate enum type will be available for your convenience, mapped from a string value.

{% tabs %}
{% tab title="Python" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```python
from mindee.v2 import ExtractionResponse
from mindee.v2.parsing.field import FieldConfidence

def handle_response(response: ExtractionResponse):
    fields = response.inference.result.fields

    confidence = fields.get_simple_field("my_simple_field").confidence

    # compare using the enum `FieldConfidence`
    is_certain = confidence == FieldConfidence.CERTAIN
    is_lte_medium = confidence <= FieldConfidence.MEDIUM
    is_gte_low = confidence >= FieldConfidence.LOW
```

{% endtab %}

{% tab title="Node.js" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```javascript
handleResponse(response) {
  const fields = response.inference.result.fields;

  const confidence = fields.getSimpleField("my_simple_field")?.confidence

  // compare using the enum `FieldConfidence`
  const isCertain = confidence === FieldConfidence.Certain;
  const isLteMedium = FieldConfidence.lessThanOrEqual(
    confidence, FieldConfidence.Medium
  );
  const isGteLow = FieldConfidence.greaterThanOrEqual(
    confidence, FieldConfidence.Low
  );
}
```

{% endtab %}

{% tab title="PHP" %}
Using the `$response` deserialized object from either the polling response or a webhook payload.

```php
use Mindee\V2\Product\Extraction\ExtractionResponse;
use Mindee\Parsing\V2\Field\FieldConfidence;

public function handleResponse(ExtractionResponse $response)
{
    $fields = $response->inference->result->fields;

    $confidence = $fields->get('my_simple_field')->confidence;

    // compare using the enum `FieldConfidence`
    $isCertain = $confidence === FieldConfidence::Certain;
    $isLteMedium = $confidence->lessThanOrEqual(FieldConfidence::Medium);
    $isGteLow = $confidence->greaterThanOrEqual(FieldConfidence::Low);
}
```

{% endtab %}

{% tab title="Ruby" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```ruby
def handle_response(response)
  fields = response.inference.result.fields

  confidence = fields["my_simple_field"].confidence

  # compare using the class `FieldConfidence`
  FieldConfidence = Mindee::Parsing::V2::Field::FieldConfidence
  
  is_certain = confidence == FieldConfidence.CERTAIN
  is_lte_medium = confidence <= FieldConfidence.MEDIUM
  is_gte_low = confidence >= FieldConfidence.LOW
end
```

{% endtab %}

{% tab title="Java" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```java
import com.mindee.v2.product.extraction.ExtractionResponse;
import com.mindee.parsing.v2.field.FieldConfidence;

public void handleResponse(ExtractionResponseresponse) {
  var fields = response.getInference().getResult().getFields();

  // choose the appropriate field type accessor method: Simple, Object, List
  FieldConfidence confidence = fields.getSimpleField("my_simple_field")
      .getConfidence();

  // compare using the enum `FieldConfidence`
  boolean isCertain = confidence === FieldConfidence.Certain;
  boolean isLteMedium = confidence.lessThanOrEqual(FieldConfidence.Medium);
  boolean isGteLow = confidence.greaterThanOrEqual(FieldConfidence.Low);
}
```

{% endtab %}

{% tab title=".NET" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```csharp
using Mindee.V2.Parsing.Inference.Field;

public void HandleResponse(ExtractionResponse response)
{
    InferenceFields fields = response.Inference.Result.Fields;

    // nullable enum since presence depends on feature activation
    FieldConfidence? confidence = fields["my_simple_field"]
        .SimpleField.Confidence;

    // compare using the enum `FieldConfidence`
    bool isCertain = confidence == FieldConfidence.Certain;
    // cast to int for relative inequalities
    bool isLteMedium = (int?)confidence <= (int)FieldConfidence.Medium;
    bool isGteLow = (int?)confidence >= (int)FieldConfidence.Low;
}
```

{% endtab %}
{% endtabs %}

### `locations`

A list of the field's locations on the document.

The data are only filled if the [Polygons (Bounding Boxes)](/extraction-models/optional-features/polygons-bounding-boxes) feature is activated.

It's possible for a single field to have multiple locations, for example when an invoice item spans two pages.

Each location has a page index and a Polygon.

Page indexes are 0-based, so the first page is `0`.

A Polygon class contains a list of Points, the specific implementation will depend on the language.

Points are listed in clockwise order, where index `0` is top left.

Point X,Y coordinates are normalized floats from 0.0 to 1.0, relative to the page dimensions.

{% tabs %}
{% tab title="Python" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```python
from mindee.v2 import ExtractionResponse

def handle_response(response: ExtractionResponse):
    fields = response.inference.result.fields
    
    locations = fields.get_simple_field("my_simple_field").locations

    # accessing the polygon
    polygon = locations[0].polygon

    # accessing points: the Polygon class extends List[Point]
    top_x = polygon[0].x

    # alternative syntax, since the Point class extends List<float>
    # top_x = polygon[0][0]

    # there are geometry functions available in the Polygon class
    center = polygon.centroid

    # accessing the page index on which the polygon is
    page_index = locations[0].page
```

{% endtab %}

{% tab title="Node.js" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```javascript
handleResponse(response) {
  const fields = response.inference.result.fields;

  const locations = fields.getSimpleField("my_simple_field")?.locations;

  // accessing the polygon
  const polygon = locations![0].polygon!;

  // accessing points: the Polygon class extends Array<Point>
  const topX = polygon[0][0];

  // there are geometry functions available in the Polygon class
  const center = polygon.getCentroid();

  // accessing the page index on which the polygon is
  const pageIndex = locations![0].page!;
}
```

{% endtab %}

{% tab title="PHP" %}
Using the `$response` deserialized object from either the polling response or a webhook payload.

```php
use Mindee\V2\Product\Extraction\ExtractionResponse;
use Mindee\Parsing\V2\Field\FieldConfidence;

public function handleResponse(ExtractionResponse $response)
{
    $fields = $response->inference->result->fields;

    $locations = $fields->get('my_simple_field')->locations;

    // accessing the polygon
    $polygon = $locations[0]->polygon;
    
    // accessing points:
    $points = $polygon->coordinates;
    $topX = $points[0]->getX();
    // alternative syntax, since the Point class extends Array<float>
    // $topX = $points[0][0]

    // there are geometry functions available in the Polygon class
    $center = $polygon->getCentroid();
    
    // accessing the page index on which the polygon is
    $pageIndex = $locations[0]->page;
}
```

{% endtab %}

{% tab title="Ruby" %}
Using the `$response` deserialized object from either the polling response or a webhook payload.

```ruby
def handle_response(response)
  fields = response.inference.result.fields

  locations = fields.get_simple_field('my_simple_field').locations

  # accessing the polygon
  polygon = locations[0].polygon

  # accessing points:
  top_x = polygon[0].x
  # alternative syntax
  # top_x = polygon[0][0]

  # there are geometry functions available in the Polygon class
  center = polygon.centroid

  # accessing the page index on which the polygon is
  page_index = locations[0].page
end
```

{% endtab %}

{% tab title="Java" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```java
import com.mindee.geometry.Point;
import com.mindee.geometry.Polygon;
import com.mindee.v2.product.extraction.ExtractionResponse;
import com.mindee.parsing.v2.field.FieldLocation;
import java.util.List;

public void handleResponse(ExtractionResponseresponse) {
  var fields = response.getInference().getResult().getFields();

  // choose the appropriate field type accessor method: Simple, Object
  List<FieldLocation> locations = fields.getSimpleField("my_simple_field")
      .getLocations();
  
  // accessing the polygon
  Polygon polygon = locations.get(0).getPolygon();

  // accessing points
  List<Point> points = polygon.getCoordinates();
  double topX = points.get(0).getX();

  // there are geometry functions available in the Polygon class
  Point center = polygon.getCentroid();

  // accessing the page index on which the polygon is
  int pageIndex = locations.get(0).getPage();
}
```

{% endtab %}

{% tab title=".NET" %}
Using the `response` deserialized object from either the polling response or a webhook payload.

```csharp
using Mindee.V2.Parsing.Inference.Field;
using Mindee.Geometry;

public void HandleResponse(ExtractionResponse response)
{
    InferenceFields fields = response.Inference.Result.Fields;

    List<FieldLocation> locations = fields["my_simple_field"]
        .SimpleField.Locations;

    // accessing the polygon
    Polygon polygon = locations.First().Polygon;

    // accessing points: the Polygon class extends List<Point>
    double topX = polygon[0].X;

    // alternative notation, since the Point class extends List<double>
    // double topX = polygon[0][0];

    // there are geometry functions available in the Polygon class
    Point center = polygon.GetCentroid();

    // accessing the page index on which the polygon is
    int pageIndex = locations.First().Page;
}
```

{% endtab %}
{% endtabs %}


# No-Code Integration

No-Code and Low-Code integration support for Extraction models.

## Overview

The Mindee API is well-suited to integrations in no-code and low-code platforms.

It's a relatively simple RESTful API so most tools and platforms can be integrated using HTTP nodes.

## Officially Supported Integrations

We're hard at work providing official integrations for Extraction models on platforms that allow it.

Currently we have support for the following platforms:

* **n8n**: [basic support for workflows](/extraction-models/no-code-integration/n8n-workflows), official support [coming soon](https://github.com/n8n-io/n8n/pull/18986)
* **Zapier**: [officially supported](/extraction-models/no-code-integration/zapier-zaps)
* **Make**: [officially supported](/extraction-models/no-code-integration/make.com-scenarios)

Don't see support for your favorite platform? [Make a feature request!](https://feedback.mindee.com/?b=682f69c9e2404756e7e68d1c)

## Generic No-Code Integration

{% hint style="warning" %}
Only use this method if your no-code platform does not have an official integration.
{% endhint %}

You'll need to use HTTP nodes in your workflow that can POST and GET a specified URL.

First, POST as form-data to the [extraction/enqueue](/integrations/api-reference/extraction-models#post-v2-products-extraction-enqueue) route:

* The Content-Type header must be set to `multipart/form-data`
* The `Authorization` header must have only your API key as the value
* The file or URL to process, one of: `file`, `file_base64`, or `url`
* The Model ID, `model_id`

If your no-code platform doesn't support the `form-data` Content-Type, you may post as `application/json`. In this case you must use `file_base64` or `url` only.

In the response to the POST, there will be a `polling_url` attribute, save its value.

Wait 3 seconds.

Loop GET requests on the `polling_url` until the response contains a `result_url` .\
**Note**: make sure to configure the node to **not** follow redirections.

Alternatively, loop on the `polling_url` until it redirects to the result.

**Important**: in all cases, wait at least 1 second between each poll.

A GET request to the `result_url` will contain the result payload.


# make.com Scenarios

Integrating Mindee in make.com scenarios.

{% hint style="info" %}
Only use the verified [**Mindee V2 app**](https://www.make.com/en/integrations/mindee).

Community apps only work for Mindee V1.
{% endhint %}

## Before You Start

When integrating Mindee using Make you'll need a file input module.

This includes, but is not limited to:

* email
* drive providers: Google Drive, MS OneDrive, DropBox, etc
* remote servers: URL/HTTP, S3, FTP, etc
* chat: Slack, Discord, etc

In the Make documentation, take a look at the section: [Working with files](https://help.make.com/working-with-files)

## Add Mindee to a Make.com Scenario

You can use the Mindee app in any make.com scenario.

When adding a module, search for "mindee" and select **Mindee V2** verified:

<figure><img src="/files/A12ckxwnDdWSEbCsACap" alt="add Mindee V2 to make.com" width="563"><figcaption></figcaption></figure>

Next, choose the "Extract Document Data" action:

<figure><img src="/files/Np1u1YiK1LCuv5qtPjNu" alt="Selecting &#x22;Enqueue and Retrieve an Inference&#x22; in make.com scenario" width="563"><figcaption></figcaption></figure>

Once you have the "Extract Document Data" module in your scenario, you'll need to connect it to one of your [Manage API Keys](/integrations/api-keys).

For this, click on the "Create a connection" button:

<figure><img src="/files/Eu8CTksHe0tFDgXJyWeq" alt="opening the Mindee V2 connection creation in make.com" width="563"><figcaption></figcaption></figure>

In the Create a connection dialog box, fill in the following information:

* A name for your connection, it should be in the format: `MindeeV2-` + your API key's name
* your [Mindee V2 API key](/integrations/api-keys#key-creation)

Finish by clicking "Save".

<figure><img src="/files/vzsqTSo75EoYfQGcj0oZ" alt="mindee v2 connection info in make.com" width="563"><figcaption></figcaption></figure>

Now you can specify which model to use. For this click on "Search Model":

<figure><img src="/files/2eT14kZKJM1yUhjAlUwM" alt="opening the search model for mindee v2 in make.com" width="563"><figcaption></figcaption></figure>

In the dialog window, you'll be able to enter a search string.

Enter in the name of the model you want to use and click "OK".

<figure><img src="/files/mpvOqOZetjIt5SOxHiXR" alt="search for Mindee v2 model in make.com" width="563"><figcaption></figcaption></figure>

If there is only one match to your search term, the model ID will be added.

If there are several matches to your search, choose the correct one from the list:

<figure><img src="/files/8UMI7oUSQfC1Bgvf8Al6" alt="choosing a Mindee v2 model in make.com" width="563"><figcaption></figcaption></figure>

{% hint style="success" %}
**That's it, you're done!** The Mindee V2 module is now ready to accept connections.
{% endhint %}

Click on "Save".

{% hint style="info" %}
**Do not fill in the "File" information in the Mindee V2 module.**

This is done automatically by connecting an appropriate module.
{% endhint %}

## Next Steps

Now that you have Mindee integrated, you can use the results in any Make action.

Some example workflows:

* Save the raw structured data (JSON) to a drive like Google Drive or MS OneDrive
* Map fields to a table format like CSV, MS Excel, or Google Sheets
* Use extracted fields to fill in:
  * CRMs like Salesforce or HubSpot
  * online documents like Notion, MS Word, or Google Docs
* etc


# n8n Workflows

Integrating Mindee in n8n workflows.

{% hint style="warning" %}
**The existing "Mindee" node in n8n is for the legacy v1 version, it will not work for v2.**

Use the instructions below to integrate n8n.

We are working on [updating the Mindee n8n node](https://github.com/n8n-io/n8n/pull/18986).
{% endhint %}

## Use a Sample n8n Workflow

Use our provided sample n8n workflow file and use it as base for creating your own workflow.

You can also import into an existing workflow, the new nodes will simply be added.

### Import the Sample n8n Workflow

First step is to download our sample workflow file below:

{% file src="/files/G4NYfqM4O3XRjafwCwLJ" %}

Next, import it into a new or existing n8n workflow:

<figure><img src="/files/pA8LwQGC4AGYrLwUqXrF" alt="importing the sample n8n workflow"><figcaption></figcaption></figure>

Once the file is imported the nodes and steps will show up.

### Configure the Authentication

Open the "Mindee V2 Post" node by double-clicking it, and create a Header Auth Credential:

<figure><img src="/files/HSs3ww6wLYSWKoTYLtV1" alt="adding mindee authentication in n8n"><figcaption></figcaption></figure>

When creating your Credential, first choose one of your [Manage API Keys](/integrations/api-keys).

Then, set the following values in the Credential dialog box:

| Name                       | Value                               |
| -------------------------- | ----------------------------------- |
| Credential Name (top left) | `Mindee v2 -` + your API key's name |
| Header Name                | `Authorization`                     |
| Header Value               | Your API key's value                |

It should look something like this when done:

<figure><img src="/files/w5SYPaftXQy8PZBEzUrK" alt="configuring mindee authentication in n8n"><figcaption></figcaption></figure>

Next, make sure the following n8n nodes are all using the Credential just created:

* "Mindee V2 GET Job"
* "Mindee V2 GET Inference"

It should be done automatically, but please open each node and make sure!

When all nodes are using the same Credential, authentication is set up correctly.

### Configure the Model ID

First copy your model's ID from the Mindee platform.

Open the "Mindee V2 Post" node, and scroll to the "Body Parameters" section.

Set the value of the "model\_id" parameter to your model's ID.

It should look something like this:

<figure><img src="/files/bnVlBy5zMR4ywzJBrL6S" alt="mindee model configuration in n8n"><figcaption></figcaption></figure>

The workflow is now configured to use the specified model.

### Test the n8n Workflow

We've included a "HTTP Request" node that will download a sample invoice file from our Github repository.

In n8n, simply click the "Execute workflow" button and follow along as the various nodes/steps are executed.

In this way make sure the workflow is able to arrive at the "Result" step. The actual contents of the result don't matter, since the sample file is there only for testing the API connection. The sample file is not meant for testing your model.

Once you're able to get a result, change the "HTTP Request" node to something actually useful to you.

This includes, but is not limited to:

* email
* drive providers: Google Drive, MS OneDrive, DropBox, etc
* remote servers: URL/HTTP, S3, FTP, etc
* chat: Slack, Discord, etc

#### Something not working as expected?

Check the [Error Handling](/integrations/problem-database) section.


# Zapier Zaps

Integrating Mindee in Zapier Zaps.

## Before You Start <a href="#add-mindee-to-a-make.com-scenario" id="add-mindee-to-a-make.com-scenario"></a>

When integrating Mindee using Zapier you'll need a trigger that exposes a File Object.

This includes, but is not limited to:

* email
* drive providers: Google Drive, MS OneDrive, DropBox, etc
* remote servers: URL/HTTP, S3, FTP, etc
* chat: Slack, Discord, etc

Once you have a trigger working, you can connect it to the Mindee action.

## Add Mindee Data Extraction to a Zap <a href="#add-mindee-to-a-make.com-scenario" id="add-mindee-to-a-make.com-scenario"></a>

You can use the [Mindee Zapier App](https://zapier.com/apps/mindee-ocr) as an action step in any Zap.

### Setup App

When adding a step, search for "*mindee*" and select **Mindee OCR**:

<figure><img src="/files/VcYpIJzcMaZrKgFO2DV2" alt="Zapier choose Mindee app" width="516"><figcaption></figcaption></figure>

You should now have the Mindee App installed.

Next choose the "Action event" by clicking on "**Document Data Extraction**":

<figure><img src="/files/19tWvy2V10kzlbZlX6vC" alt="Zapier choose Mindee data extraction action" width="563"><figcaption></figcaption></figure>

Next, you'll need to connect the App to one of your [Manage API Keys](/integrations/api-keys).

If you already have a Mindee connection configured in Zapier, you'll have the option to "select" a connection in the "Account" section.

Otherwise, click on the "Sign in" button in the "Account" section.

Then, in the connection dialog box, fill in the following information:

* A name for your connection, for example your API key's name
* your [Mindee V2 API key](/integrations/api-keys#key-creation)

Finish by clicking on "Yes, Continue to Mindee OCR":

<figure><img src="/files/o8UIw1t6yZv9SETWSNis" alt="Zapier create connection Mindee API key" width="476"><figcaption></figcaption></figure>

{% hint style="info" %}
You'll be able to re-use this connection in other Zaps that include the Mindee App.
{% endhint %}

Your Step should now look something like this:

<figure><img src="/files/4f0NwNKyYpOA6KajdSiE" alt="" width="313"><figcaption></figcaption></figure>

Next, to configure the action, click on "Continue" at the bottom of the screen.

### Configure the Data Extraction Action

First, click on "Model to Use", you'll then be given a list of models available in your organization\
Click the one that you want to use:

<figure><img src="/files/SL1NNxo5ogP7IEwD4NuS" alt="" width="563"><figcaption></figcaption></figure>

For the "File to Send", you'll need to connect to a valid file input.

Normally it will be a field called "**File (Exists but not shown)**".\
As an example, here is what it looks like when connecting to a Drive folder:

<figure><img src="/files/RU4cnWgsE9T1FlFnCG3f" alt="" width="374"><figcaption></figcaption></figure>

You can also set various options:

* [Raw Text (Full OCR)](/extraction-models/optional-features/raw-text-full-ocr)\
  Add the full text content of your documents to the API response.
* [Continuous Learning (RAG)](/extraction-models/optional-features/improving-accuracy)\
  Enhance extraction accuracy with Retrieval-Augmented Generation using your own documents.
* [Polygons (Bounding Boxes)](/extraction-models/optional-features/polygons-bounding-boxes)\
  Add the polygon coordinates of each extracted field to the API response.
* [Confidence Score and Accuracy Boost](/extraction-models/optional-features/automation-confidence-score)\
  🚀 Boost the precision and accuracy of all extractions.\
  Add a confidence score to each extracted field.
* "Polling Timeout" - The maximum number of seconds to attempt retrieving results, before stopping with an error. Increase this value if you are sending documents with many pages and are consistently getting timeout errors.

Next, to test the step, click on "Continue" at the bottom of the screen.

### Test the Step

If everything in the previous steps was configured correctly, you should now have something that looks like this:

<figure><img src="/files/Cug9WcPb4okXlKnUZqdN" alt="" width="311"><figcaption></figcaption></figure>

After clicking on "Test step", the file will be sent to Mindee for processing.

When complete, it should look like this:

<figure><img src="/files/SToCGNt7Obl8TJAACr9S" alt="" width="308"><figcaption></figcaption></figure>

Congrats, everything works!

## Next Steps

Now that you have Mindee integrated, you can use the results in any Step of your Zap.

Some example workflows:

* Save the raw structured data (JSON) to a drive like Google Drive or MS OneDrive
* Map fields to a table format like CSV, MS Excel, or Google Sheets
* Use extracted fields to fill in:
  * CRMs like Salesforce or HubSpot
  * online documents like Notion, MS Word, or Google Docs
* etc


# Split Model Overview

Automatically breaking a multi-page source file into separate documents and associate a class to each one.

## Use Cases

Process a bundle of different documents sent in the same file. The result has both the page range and the class for each document identified, allowing for complex workflows.

Some common examples, where a single PDF contains:

* Several different invoices
* A mix of invoices, receipts, and bank statements
* The person's driver license, vehicle registration, insurance
* Front and back of an ID card, each on a separate page
* The same type of document, but from different regions or languages

A file sent to the Split Model may have any number of pages, [within limits](/integrations/technical-limitations#file-limits).

## Create a Split Model

Split models are always custom, there are no templates available in the Catalog.\
This keeps Split models flexible for different documents and workflows.

Each Split model gets its own unique model ID when you create it.

1. To create a Split model, you need to click on **Models**, and then on **Create your document AI model**.
2. Scroll to the **Document Utilities** section, click on **Split.**
3. A pop-up will appear, allowing you to enter the classes you want. Each class corresponds to a document type possibly present in the documents you want to process.\
   \
   For example, if the files you are processing contain invoices, receipts, and driving licenses, set the classes as: `INVOICE`, `RECEIPT`, `DRIVER_LICENSE`.

{% hint style="info" icon="lightbulb" %}
Add the class `OTHER` if you need the model to identify documents that are not one of the explicitly defined classes.
{% endhint %}

<figure><img src="/files/bQzEcFgC6ocOI6lDCgLF" alt="" width="375"><figcaption></figcaption></figure>

4. Once ready, click on **Create Utility** to create your custom Split Model.\
   This step will also generate the model's unique ID.
5. You can now use the **Live Test** tab to process documents, and the **Utility Configuration** to update your classes.<br>

Your utility is now available in your **Models** tab:

<figure><img src="/files/lEz0jYFEDhSzyIhXIy96" alt=""><figcaption></figcaption></figure>

Here is a step-by-step tutorial that shows you how to properly create a Split Utility :<br>

{% @supademo/embed url="<https://app.supademo.com/demo/cmls4fbup1r1611891nvrrlyw>" demoId="cmls4fbup1r1611891nvrrlyw" %}

## Technical Considerations

Models requires at least two classes defined.

Class names may be in most languages and writing systems, but cannot exceed 128 characters.

Having more than a few dozen classes will yield unexpected results, and is not recommended. If you need classes based on vendor/customer names, product codes, phone numbers, etc, you should use a text field in an [Extraction model](/extraction-models/extraction-models-overview) instead.

{% hint style="info" %}
Class names will be returned **exactly** as defined on the platform in the return, spaces and capitalization included.

If the class names are changed on the platform, the change in the API return will be **immediate** for all new files sent.
{% endhint %}

## Integration

Once your Split model is created and tested, integration documentation is provided in the "Documentation" page, or here: [Split Quick Start](/split-models/sdk-integration/split-quick-start).


# Extraction Model Chaining

Extract data from detected split ranges.

Use Split ranges to automatically extract document data, allowing for several different extractions on a multi-page file.

Note: Split Models also work on single-page files, in this case the range will always be exactly one page. If you need multiple extraction results on the same page, take a look at [Crop Models](/crop-models/crop) instead.

## Extraction Set Up

### At Model Creation

When creating your Split model, you'll be adding document classes in the creation window.

For each document class, you can set one of your [Extraction Models](/extraction-models/extraction-models-overview) for chaining. The Extraction Model must exist prior to the Split Model creation.

Use the search field to filter available extraction models.

<figure><img src="/files/kiD2AxsdWZ0f0XQo5IOz" alt="Configuring extraction model chaining on Split creation" width="563"><figcaption></figcaption></figure>

### After Model Creation

This works exactly like when creating at model creation.

Simply go to your Split model's "Utility Configuration" page and adjust as needed.

You can add new classes, remove classes, and change Extraction Models.

<figure><img src="/files/7r9Z6Ls8OJ1bqcIaWOkE" alt="Configuring extraction model chaining on Split after creation" width="563"><figcaption></figcaption></figure>

### Selectively Extracting

If a detected class has no linked Extraction Model, no extraction runs for that crop.

This allows selectively extracting some sections of the file while ignoring others.

Let's say you receive large multi-page PDFs from your users, where each PDF is a bundle of different scanned documents: plane tickets, travel receipts, driver license, and passport.

You need only the passports.

In your Split configuration, add a `passport` class and an `other`  class, and only link an extraction model to the `passport` class.

All split ranges will get classified, but only those linked to an Extraction Model will have extraction results.

{% hint style="info" icon="lightbulb" %}
Add an `other` class in addition to classes of interest to you, unless all potential document types are known in advance.
{% endhint %}

#### Removing Unused Pages

It's also possible to remove pages that are never used in the Extraction. For example to remove terms and conditions from invoices, set up the classes `terms_and_conditions_page` and `invoice_page` , and only link an Extraction Model to the `invoice_page` class.

### Token Usage With Chaining

You only consume tokens once, there is no double charge when chaining to an Extraction model.

The possible credit consumption scenarios are as follows:

* extraction model not chained ⇒ token usage of initial model (Split, Crop, Classification)
* extraction model chained ⇒ token usage of the configured extraction model, including any activated [optional features](/extraction-models/optional-features).

In other words, the initial model is provided at no additional cost to you if it triggers additional processing using an Extraction model.

All token consumption is per page as is standard for our models.

## Access Extraction Results

When an Extraction model is linked, each Split Range with the detected class will contain an Extraction Response object, which is identical when making an Extraction request for a single document.

Check the [Split Result](/split-models/sdk-integration/split-result) section for more details.


# SDK Integration

Integrate a Split model using the Mindee SDKs.

Use the SDKs to send multi-page files to a Split model and process split results.

This section helps you choose the right starting point and move through the full integration flow.

### Choose your path

Start with the page that matches your next step:

* [Split Quick Start](/split-models/sdk-integration/split-quick-start) ⇒ install a client library, send a file, and get your first result.
* [Split Result](/split-models/sdk-integration/split-result) ⇒ access split ranges, document types, and optional chained extraction results.
* [Split Model Overview](/split-models/split) ⇒ understand how Split models group pages into logical documents.

### Typical SDK workflow

{% stepper %}
{% step %}

#### Create Your Model

[Create your Split Model](/split-models/split#create-a-split-model) on the Mindee platform.

Upload some samples to the [Live Test](/models/live-test) to validate the model.
{% endstep %}

{% step %}

#### Install and authenticate

Install the client library for your language.

Prepare your API key and initialize the client.
{% endstep %}

{% step %}

#### Send a file

Set the Split model's ID and send a file or URL for processing.

Start with polling unless you already use webhooks.
{% endstep %}

{% step %}

#### Process the split results

Read the list of detected documents from the response.

Use each split's document type and page range in your workflow.
{% endstep %}
{% endstepper %}

### Before you start

Have these ready:

* your API key
* your Split Model's unique ID
* a multi-page sample file for testing
* a language choice for the SDK

### What is specific to Split models

Split responses contain a list of logical documents found in the source file.

Each split includes a document type and a 0-based page range.

Document types are returned exactly as configured on the platform.

If extraction chaining is enabled, each split can also include an extraction response.

### Shared SDK building blocks

Integration builds on the same client library concepts used across all model types.

* [Client Libraries / SDKs](/integrations/client-libraries-sdk)
* [Client Configuration](/integrations/client-libraries-sdk/configure-the-client)
* [Send a File or URL](/integrations/client-libraries-sdk/send-a-file-or-url)
* [Response Processing](/integrations/client-libraries-sdk/process-the-response)
* [Webhook Results](/integrations/webhooks)
* [Error Handling](/integrations/problem-database)


# Split Quick Start

Integrate a Split Model.

## Install the Client Library

Install the Mindee SDK for your language or framework of choice

{% tabs %}
{% tab title="Python" %}
Requires Python ≥ 3.9. Python ≥ 3.11 is recommended.

Simply install the [PyPi package](https://pypi.org/project/mindee/) using `pip`:

```sh
pip install -U mindee~=5.2
```

{% endtab %}

{% tab title="Node.js" %}
Requires Node.js ≥ 20.1. Node.js ≥ 22 is recommended.

Simply install the [NPM package](https://www.npmjs.com/package/mindee):

```sh
npm install mindee@^5.5.0
```

{% endtab %}

{% tab title="PHP" %}
Requires PHP ≥ 8.1. PHP ≥ 8.3 is recommended.

Simply install the [Packagist package](https://packagist.org/packages/mindee/mindee) using [composer](https://getcomposer.org/):

```sh
php composer.phar require "mindee/mindee:>=3.0"
```

{% endtab %}

{% tab title="Ruby" %}
Requires Ruby ≥ 3.2.

Simply install the [gem](https://rubygems.org/gems/mindee) using:

```shell
gem install mindee -v '~> 5.2'
```

{% endtab %}

{% tab title="Java" %}
Requires Java ≥ 11. Java ≥ 17 is recommended.

Group ID: `com.mindee.sdk`\
Artifact ID: `mindee-api-java`\
Version: `5.2.0` or greater

There are various installation methods, Maven, Gradle, etc:

[Installation Details](https://central.sonatype.com/artifact/com.mindee.sdk/mindee-api-java)
{% endtab %}

{% tab title=".NET" %}
.NET ≥ 8.0 is recommended.

Simply install the [NuGet package](https://www.nuget.org/packages/Mindee) using `dotnet add`:

```sh
dotnet add package Mindee --version 4.4
```

{% endtab %}
{% endtabs %}

Don't see support for your favorite language or framework? [Make a feature request!](https://feedback.mindee.com/?b=682f69c9e2404756e7e68d1c)

## Send a File and Poll

Use the client library to send the file to your Split Model and return the result.

{% tabs %}
{% tab title="Python" %}
Requires Python ≥ 3.9. Python ≥ 3.11 is recommended.\
Requires the [Mindee Python SDK](https://pypi.org/project/mindee/) version **4.36.0** or greater.

{% code lineNumbers="true" %}

```python
from mindee import PathInput
from mindee.v2 import (
    Client,
    SplitParameters,
    SplitResponse,
)

input_path = "/path/to/the/file.ext"
api_key = "MY_API_KEY"
model_id = "MY_MODEL_ID"

# Init a new client
mindee_client = Client(api_key)

# Set Split parameters
model_params = SplitParameters(
    # ID of the model, required.
    model_id=model_id,
)

# Load a file from disk
input_source = PathInput(input_path)

# Send for processing using polling
response = mindee_client.enqueue_and_get_result(
    SplitResponse,
    input_source,
    model_params,
)

# Print a brief summary of the parsed data
print(response.inference)

# Access the split result
splits: list = response.inference.result.splits
```

{% endcode %}

Also take a look at the [Split Result](https://docs.mindee.com/split-models/sdk-integration/split-result) documentation.
{% endtab %}

{% tab title="Node.js" %}
Requires Node.js ≥ 20.1. Node.js ≥ 22 is recommended.\
Requires the [Mindee Node.js SDK](https://www.npmjs.com/package/mindee/) version **5.5.0** or greater.

{% code lineNumbers="true" %}

```javascript
import * as mindee from "mindee";
// If you're on CommonJS:
// const mindee = require("mindee");

const apiKey = "MY_API_KEY";
const filePath = "/path/to/the/file.ext";
const modelId = "MY_MODEL_ID";

// Init a new client
const mindeeClient = new mindee.Client(
  { apiKey: apiKey }
);

// Set Split parameters
const modelParams = {
  modelId: modelId,
};

// Load a file from disk
const inputSource = new mindee.PathInput({ inputPath: filePath });

// Send for processing
const response = await mindeeClient.enqueueAndGetResult(
  mindee.product.Split,
  inputSource,
  modelParams,
);

// print a string summary
console.log(response.inference.toString());

// Access the result splits
const crops = response.inference.result.splits;
```

{% endcode %}

Also take a look at the [Split Result](https://docs.mindee.com/split-models/sdk-integration/split-result) documentation.
{% endtab %}

{% tab title="PHP" %}
Requires PHP ≥ 8.1. PHP ≥ 8.3 is recommended.\
Requires the [Mindee PHP SDK](https://packagist.org/packages/mindee/mindee) version **3.0.0** or greater.

{% code lineNumbers="true" %}

```php
<?php

use Mindee\Input\PathInput;
use Mindee\V2\Client;
use Mindee\V2\Product\Split\Params\SplitParameters;
use Mindee\V2\Product\Split\SplitResponse;

$apiKey = "MY_API_KEY";
$modelId = "MY_MODEL_ID";
$filePath = "/path/to/the/file.ext";

// Init a new client
$mindeeClient = new Client($apiKey);

// Set Split parameters
$modelParams = new SplitParameters(
    // ID of the model, required.
    $modelId,
);

// Load a file from disk
$inputSource = new PathInput($filePath);

// Send for processing using polling
$response = $mindeeClient->enqueueAndGetResult(
    SplitResponse::class,
    $inputSource,
    $modelParams
);

// Print a summary of the response
echo strval($response->inference);

// Access the split results
$splits = $response->inference->result->splits;
```

{% endcode %}

Also take a look at the [Split Result](https://docs.mindee.com/split-models/sdk-integration/split-result) documentation.
{% endtab %}

{% tab title="Ruby" %}
Requires Ruby ≥ 3.2.\
Requires the [Mindee Ruby SDK](https://rubygems.org/gems/mindee) version **5.2.1** or greater.

{% code lineNumbers="true" %}

```ruby
require 'mindee'
require 'mindee/v2/product'

input_path = '/path/to/the/file.ext'
api_key = 'MY_API_KEY'
model_id = 'MY_MODEL_ID'

# Init a new client
mindee_client = Mindee::V2::Client.new(api_key: api_key)

# Set Split parameters
model_params = {
    # ID of the model, required.
    model_id: model_id,
}

# Load a file from disk
input_source = Mindee::Input::Source::PathInputSource.new(input_path)

# Send for processing
response = mindee_client.enqueue_and_get_result(
    Mindee::V2::Product::Split::Split,
    input_source,
    model_params
)

# Access the result splits
puts response.inference.result.splits
```

{% endcode %}

Also take a look at the [Split Result](https://docs.mindee.com/split-models/sdk-integration/split-result) documentation.
{% endtab %}

{% tab title="Java" %}
Requires Java ≥ 11. Java ≥ 17 is recommended.\
Requires the [Mindee Java SDK](https://central.sonatype.com/artifact/com.mindee.sdk/mindee-api-java) version **5.1.0** or greater.

{% code lineNumbers="true" %}

```java
import com.mindee.input.LocalInputSource;
import com.mindee.v2.MindeeClient;
import com.mindee.v2.product.split.SplitResponse;
import com.mindee.v2.product.split.params.SplitParameters;
import java.io.IOException;

public class SimpleMindeeClientV2 {

  public static void main(String[] args)
      throws IOException, InterruptedException
  {
    String apiKey = "MY_API_KEY";
    String modelId = "MY_MODEL_ID";
    String filePath = "/path/to/the/file.ext";

    // Init a new client
    var mindeeClient = new MindeeClient(apiKey);

    // Set Split parameters
    var modelParams = SplitParameters
        // ID of the model, required.
        .builder(modelId)
        .build();

    // Load a file from disk
    var inputSource = new LocalInputSource(filePath);

    // Send for processing using polling
    SplitResponse response = mindeeClient.enqueueAndGetResult(
        SplitResponse.class,
        inputSource,
        modelParams
    );

    // Print a summary of the response
    System.out.println(response.getInference().toString());

    // Access the split result
    var splits = response.getInference().getResult().getSplits();
  }
}
```

{% endcode %}

Also take a look at the [Split Result](https://docs.mindee.com/split-models/sdk-integration/split-result) documentation.
{% endtab %}

{% tab title=".NET" %}
.NET ≥ 8.0 is recommended.\
Requires the [Mindee .NET SDK](https://www.nuget.org/packages/Mindee) version **4.3.0** or greater.

{% code lineNumbers="true" %}

```csharp
using Mindee;
using Mindee.Input;
using Mindee.V2;
using Mindee.V2.Product.Split;
using Mindee.V2.Product.Split.Params;

string filePath = "/path/to/the/file.ext";
string apiKey = "MY_API_KEY";
string modelId = "MY_MODEL_ID";

// Construct a new client
Client mindeeClient = new Client(apiKey);

// Set Split parameters
var modelParams = new SplitParameters(
    modelId: modelId
);

// Load a file from disk
var inputSource = new LocalInputSource(filePath);

// Upload the file
var response = await mindeeClient.EnqueueAndGetResultAsync<SplitResponse>(
    inputSource, modelParams);

// Print a summary of the response
System.Console.WriteLine(response.Inference.ToString());

// Access the split range results
var splitRanges = response.Inference.Result.Splits;
```

{% endcode %}

Also take a look at the [Split Result](https://docs.mindee.com/split-models/sdk-integration/split-result) documentation.
{% endtab %}
{% endtabs %}


# Split Result

Reference documentation on processing a Split result using the Mindee SDKs.

You'll need to have a response object as described in the [Response Processing](/integrations/client-libraries-sdk/process-the-response) section.

## Access Split Ranges

A `SplitRange` describes one logical sub-document identified within the source file.

### `SplitRange` Attributes

Each `SplitRange` instance will have these properties.

#### Document Type

The document category assigned to the sub-document. It is always filled.

It is returned as a string and is identical to the value entered on the Mindee platform — case, spaces, and punctuation included.

#### Page Range

A two-element array of **0-based** page indexes, where the first integer indicates the start page and the second integer indicates the end page. It is always filled.

A page range value of `(0,2)` would mean from the **first** page to the **third** page.

#### Extraction Response

Optional extraction response associated with the split. This is only filled if extraction chaining is activated for the model.

### Iterate Over Split Ranges

You'll usually want to iterate over all split ranges, since the number of ranges is dependent on the document.

{% tabs %}
{% tab title="Python" %}

```python
from mindee import SplitResponse

def handle_response(response: SplitResponse) -> None:
    splits = response.inference.result.splits

    for split in splits:
        # 0-based [start_page, end_page] range
        page_range = split.page_range

        print(f"Detected type: {split.document_type}")
        print(f"On pages: {page_range[0]} - {page_range[1]}")

        # Optional extraction response, present if extraction chaining was requested
        extraction_response = split.extraction_response

        if extraction_response is not None:
            # Access extracted fields from the split's inference result
            fields = extraction_response.inference.result.fields
            print(f"Extraction fields: {fields}")
```

{% endtab %}

{% tab title="Node.js" %}

```javascript
handleResponse(response) {
  const splits = response.inference.result.splits;

  for (const split of splits) {
    // 0-based [startPage, endPage] range
    const pageRange = split.pageRange;

    // Document type identified for this split
    const documentType = split.documentType;

    console.log(`Detected type: ${documentType}`);
    console.log(`On pages: ${pageRange[0]} - ${pageRange[1]}`);
    
    // Optional extraction response, present if extraction chaining was requested
    const extractionResponse = split.extractionResponse;

    if (extractionResponse) {
      // Access extracted fields from the split's inference result
      const fields = extractionResponse.inference.result.fields;
      console.log("Extraction fields:", fields.toString());
    }
  }
}
```

{% endtab %}

{% tab title="PHP" %}

```php
use Mindee\V2\Product\Split\SplitResponse;

public function handleResponse(SplitResponse $response)
{
    $splits = $response->inference->result->splits;

    foreach ($splits as $split) {
        // 0-based [startPage, endPage] range
        $pageRange = $split->pageRange;

        // Document type identified for this split
        $documentType = $split->documentType;

        echo "Detected type: $documentType\n";
        echo "On pages: {$pageRange[0]} - {$pageRange[1]}\n";

        // Optional extraction response, present if extraction chaining was requested
        $extractionResponse = $split->extractionResponse;

        if ($extractionResponse !== null) {
            // Access extracted fields from the split's inference result
            $fields = $extractionResponse->inference->result->fields;
            echo "Extraction fields: " . $fields->toString() . "\n";
        }
    }
}
```

{% endtab %}

{% tab title="Ruby" %}

```ruby
def handle_response(response)
  splits = response.inference.result.splits

  splits.each do |split|
    # 0-based [start_page, end_page] range
    page_range = split.page_range

    # Document type identified for this split
    document_type = split.document_type

    puts "Detected type: #{document_type}"
    puts "On pages: #{page_range[0]} - #{page_range[1]}"

    # Optional extraction response, present if extraction chaining was requested
    extraction_response = split.extraction_response

    if extraction_response
      # Access extracted fields from the split's inference result
      fields = extraction_response.inference.result.fields
      puts "Extraction fields: #{fields}"
    end
  end
```

{% endtab %}

{% tab title="Java" %}

```java
import com.mindee.v2.product.split.SplitResponse;
import com.mindee.v2.product.split.SplitRange;

public void handleResponse(SplitResponse response) {
  var splits = response.getInference().getResult().getSplits();

  for (SplitRange split : splits) {
    // 0-based [startPage, endPage] range
    var pageRange = split.getPageRange();

    // Document type identified for this split
    String documentType = split.getDocumentType();

    System.out.println("Detected type: " + documentType);
    System.out.println(
      "On pages: " + pageRange.get(0) + " - " + pageRange.get(1)
    );

    // Optional extraction response, present if extraction chaining was requested
    var extractionResponse = split.getExtractionResponse();

    if (extractionResponse != null) {
      // Access extracted fields from the split's inference result
      var fields = extractionResponse.getInference().getResult().getFields();
      System.out.println("Extraction fields: " + fields.toString());
    }
  }
}
```

{% endtab %}

{% tab title=".NET" %}

```csharp
using System;
using Mindee.V2.Product.Split;

public void HandleResponse(SplitResponse response)
{
    var splits = response.Inference.Result.Splits;

    foreach (SplitRange split in splits)
    {
        // 0-based [startPage, endPage] range
        var pageRange = split.PageRange;

        // Document type identified for this split
        string documentType = split.DocumentType;

        Console.WriteLine($"Detected type: {documentType}");
        Console.WriteLine($"On pages: {pageRange[0]} - {pageRange[1]}");

        // Optional extraction response, present if extraction chaining was requested
        var extractionResponse = split.ExtractionResponse;

        if (extractionResponse != null)
        {
            // Access extracted fields from the split's inference result
            var fields = extractionResponse.Inference.Result.Fields;
            Console.WriteLine($"Extraction fields: {fields}");
        }
    }
}
```

{% endtab %}
{% endtabs %}

Chained Extraction results will be an `ExtractionResponse` object. This object is identical to the response of an Extraction Model.

Refer to the [Extraction Result](/extraction-models/sdk-integration/extraction-result) section for details on processing the Extraction.

## Extract Split Ranges From the Input File

The SDKs provide support for extracting each split range as separate PDF files.

{% tabs %}
{% tab title="Python" %}

```python
from mindee import SplitResponse

def handle_response(response: SplitResponse, input_source) -> None:
    splits = response.inference.result.splits

    for split in splits:
      extracted_split = split.extract_from_input_source(input_source)

```

{% endtab %}

{% tab title="Node.js" %}

<pre class="language-javascript"><code class="lang-javascript"><strong>async handleResponse(response, inputSource) {
</strong>  const splits = response.inference.result.splits;

  for (const split of splits) {
    const extractedSplit = await split.extractFromInputSource(inputSource);
  }
}
</code></pre>

{% endtab %}

{% tab title="PHP" %}

```php
use Mindee\V2\Product\Split\SplitResponse;

public function handleResponse(SplitResponse $response, $inputSource)
{
    $splits = $response->inference->result->splits;

    foreach ($splits as $split) {
        $extractedSplit = $split->extractFromInputSource($inputSource)
    }
}
```

{% endtab %}

{% tab title="Ruby" %}

```ruby
def handle_response(response)
  splits = response.inference.result.splits

  splits.each do |split|
    extracted_split = split.extract_from_input_source(input_source)
  end
```

{% endtab %}

{% tab title="Java" %}

```java
import com.mindee.input.LocalInputSource;
import com.mindee.v2.product.split.SplitResponse;
import com.mindee.v2.product.split.SplitRange;

public void handleResponse(
  SplitResponse response,
  LocalInputSource inputSource
) {
  var splits = response.getInference().getResult().getSplits();

  for (SplitRange split : splits) {
    var extractedSplit = split.extractFromInputSource(inputSource);
  }
}
```

{% endtab %}

{% tab title=".NET" %}

```csharp
using System;
using Mindee.Input;
using Mindee.V2.Product.Split;

public void HandleResponse(SplitResponse response, LocalInputSource inputSource)
{
    var splits = response.Inference.Result.Splits;

    foreach (SplitRange split in splits)
    {
        var extractedSplit = split.ExtractFromInputSource(inputSource);
    }
}
```

{% endtab %}
{% endtabs %}

{% hint style="info" %}
**Only the methods documented here are production-ready**.

This is because we are still improving the SDKs, hang tight!
{% endhint %}


# No-Code Integration

No-Code and Low-Code integration support for Split models.

## Officially Supported Integrations

We're hard at work providing official integrations for Split models on platforms that allow it.

Currently we have support for the following platforms:

* **n8n**: official support [coming soon](https://github.com/n8n-io/n8n/pull/18986)
* **Zapier**: officially supported
* **Make**: officially supported

Don't see support for your favorite platform? [Make a feature request!](https://feedback.mindee.com/?b=682f69c9e2404756e7e68d1c)

{% hint style="info" %}
**Usage is very similar to** [**Extraction models**](/extraction-models/no-code-integration)**.**

Complete documentation in progress...
{% endhint %}


# Crop Model Overview

Automatically identify the borders of documents on each page, and match one to a category.

## Use Cases

Process different documents sent on the same page (usually a photo). The result has both the location (page and coordinates) and the class for each document identified, allowing for complex workflows.

A file sent to the Crop Model may have any number of pages, [within limits](/integrations/technical-limitations#file-limits). Use the page index in the results to identify on which page the document was found.

Some common examples:

* Single photo of a bunch of receipts on the table
* Front and back of an ID card, each on the same PDF page
* Remove the background from all documents in a multi-page PDF

## Create a Crop Model

Crop models are always custom, there are no templates available in the Catalog.\
This keeps Crop models flexible for different documents and workflows.

Each Crop model gets its own unique model ID when you create it.

1. To create a Crop model, click on **Models**, and then click on **Create your document AI model**.
2. Scroll to the **Document Utilities** section, click on **Crop.**
3. A pop-up will appear, allowing you to enter the classes you want. Each class corresponds to a document type possibly present in the pages you want to process.\
   \
   For example, if the files you are processing contain ID cards and passports, set the classes as:\
   &#x20;`ID Card Front`, `ID Card Back`, `Passport`.

{% hint style="info" icon="lightbulb" %}
Add the class `OTHER` if you need the model to identify documents that are not one of the explicitly defined classes.
{% endhint %}

<figure><img src="/files/mqlmQ4dW25gx3mtyChv8" alt="" width="375"><figcaption></figcaption></figure>

4. Once ready, click on **Create Utility** to create your custom Crop Model.\
   This step will also generate the model's unique ID.
5. You can now use the **Live Test** tab to process documents, and the **Utility Configuration** to update your classes.<br>

Your utility is now available in your **Models** tab:

<figure><img src="/files/vPKFNVaYq2bfTtlCByhi" alt=""><figcaption></figcaption></figure>

Here is a step-by-step tutorial that shows you how to properly create a Crop Utility :

{% @supademo/embed url="<https://app.supademo.com/demo/cmlrtp7mk0u531189vkn2azqr>" demoId="cmlrtp7mk0u531189vkn2azqr" %}

## Technical Considerations

Models requires at least two classes defined.

Class names may be in most languages and writing systems, but cannot exceed 128 characters.

Having more than a few dozen classes will yield unexpected results, and is not recommended. If you need classes based on vendor/customer names, product codes, phone numbers, etc, you should use a text field in an [Extraction model](/extraction-models/extraction-models-overview) instead.

{% hint style="info" %}
Class names will be returned **exactly** as defined on the platform in the return, spaces and capitalization included.

If the class names are changed on the platform, the change in the API return will be **immediate** for all new files sent.
{% endhint %}

## Integration

Once your Crop model is created and tested, integration documentation is provided in the "Documentation" page.


# Extraction Model Chaining

Extract data from detected crops.

Use Crop to automatically extract document data, meaning that several different extractions can be made for a single image.

Note: Crop Models also work on multi-page files, with potentially multiple crop items for each page. However each crop item is limited to an area on a single page. If you need a single extraction result covering multiple pages, take a look at [Split Models](/split-models/split) instead.

## Extraction Set Up

### At Model Creation

When creating your Crop model, you'll be adding document classes in the creation window.

For each document class, you can link one of your [Extraction Models](/extraction-models/extraction-models-overview) for chaining. The Extraction Model must exist prior to the Crop Model creation.

Use the search field to filter available extraction models.

### After Model Creation

This works exactly like when creating at model creation.

Simply go to your Crop Model's "Utility Configuration" page and adjust as needed.

You can add new classes, remove classes, and change Extraction Models.

### Selectively Extracting

If a detected class has no linked Extraction Model, no extraction runs for that crop.

This allows selectively extracting some parts while ignoring others.

Let's say your users upload images of their trip documents on their hotel table, typically but not limited to: plane tickets, travel receipts, driver license, and passport.

You need only the passports.

In your Crop configuration, add a `passport` class and an `other`  class, and only link an extraction model to the `passport` class.

All crop items will get classified, but only those linked to an Extraction Model will have extraction results.

{% hint style="info" icon="lightbulb" %}
Add an `other` class in addition to classes of interest to you, unless all potential document types are known in advance.
{% endhint %}

### Token Usage With Chaining

You only consume tokens once, there is no double charge when chaining to an Extraction model.

The possible credit consumption scenarios are as follows:

* extraction model not chained ⇒ token usage of initial model (Split, Crop, Classification)
* extraction model chained ⇒ token usage of the configured extraction model, including any activated [optional features](/extraction-models/optional-features).

In other words, the initial model is provided at no additional cost to you if it triggers additional processing using an Extraction model.

All token consumption is per page as is standard for our models.

## Access Extraction Results

When an Extraction Model is linked, each detected crop item with that class contains an Extraction Response object.

That object is the same as the response returned by a standalone Extraction request.

Because Crop works at object level, several crop items on the same page can each contain their own extraction result.

Check [Crop Result](/crop-models/sdk-integration/crop-result) for details on accessing crop items and their extraction responses.


# SDK Integration

Integrate a Crop model using the Mindee SDKs.

Use the SDKs to send files to a Crop model and process crop results.

This section helps you choose the right starting point and move through the full integration flow.

### Choose your path

Start with the page that matches your next step:

* [Crop Quick Start](/crop-models/sdk-integration/crop-quick-start) ⇒ install a client library, send a file, and get your first result.
* [Crop Result](/crop-models/sdk-integration/crop-result) ⇒ access detected items, object types, and optional chained extraction results.
* [Crop Model Overview](/crop-models/crop) ⇒ understand how Crop models detect documents or objects on a page.

### Typical SDK workflow

{% stepper %}
{% step %}

#### Create Your Model

[Create your Crop Model](/crop-models/crop#create-a-crop-model) on the Mindee platform.

Upload some samples to the [Live Test](/models/live-test) to validate the model.
{% endstep %}

{% step %}

#### Install and authenticate

Install the client library for your language.

Prepare your API key and initialize the client.
{% endstep %}

{% step %}

#### Send a file

Set the Crop model's ID and send a file or URL for processing.

Start with polling unless you already use webhooks.
{% endstep %}

{% step %}

#### Process the crop results

Read the list of detected items from the response.

Use each item's object type, page, and polygon in your workflow.
{% endstep %}
{% endstepper %}

### Before you start

Have these ready:

* your API key
* your Crop Model's unique ID
* a sample file for testing
* a language choice for the SDK

### What is specific to Crop models

Crop responses contain a list of detected items found in the source file.

Each item includes an object type and a location.

The location contains a 0-based page index and polygon coordinates.

Object types are returned exactly as configured on the platform.

If extraction chaining is enabled, each crop item can also include an extraction response.

### Shared SDK building blocks

Integration builds on the same client library concepts used across all model types.

* [Client Libraries / SDKs](/integrations/client-libraries-sdk)
* [Client Configuration](/integrations/client-libraries-sdk/configure-the-client)
* [Send a File or URL](/integrations/client-libraries-sdk/send-a-file-or-url)
* [Response Processing](/integrations/client-libraries-sdk/process-the-response)
* [Webhook Results](/integrations/webhooks)
* [Error Handling](/integrations/problem-database)


# Crop Quick Start

Integrate a Crop Model.

## Install the Client Library

Install the Mindee SDK for your language or framework of choice

{% tabs %}
{% tab title="Python" %}
Requires Python ≥ 3.9. Python ≥ 3.11 is recommended.

Simply install the [PyPi package](https://pypi.org/project/mindee/) using `pip`:

```sh
pip install -U mindee~=5.2
```

{% endtab %}

{% tab title="Node.js" %}
Requires Node.js ≥ 20.1. Node.js ≥ 22 is recommended.

Simply install the [NPM package](https://www.npmjs.com/package/mindee):

```sh
npm install mindee@^5.5.0
```

{% endtab %}

{% tab title="PHP" %}
Requires PHP ≥ 8.1. PHP ≥ 8.3 is recommended.

Simply install the [Packagist package](https://packagist.org/packages/mindee/mindee) using [composer](https://getcomposer.org/):

```sh
php composer.phar require "mindee/mindee:>=3.0"
```

{% endtab %}

{% tab title="Ruby" %}
Requires Ruby ≥ 3.2.

Simply install the [gem](https://rubygems.org/gems/mindee) using:

```shell
gem install mindee -v '~> 5.2'
```

{% endtab %}

{% tab title="Java" %}
Requires Java ≥ 11. Java ≥ 17 is recommended.

Group ID: `com.mindee.sdk`\
Artifact ID: `mindee-api-java`\
Version: `5.2.0` or greater

There are various installation methods, Maven, Gradle, etc:

[Installation Details](https://central.sonatype.com/artifact/com.mindee.sdk/mindee-api-java)
{% endtab %}

{% tab title=".NET" %}
.NET ≥ 8.0 is recommended.

Simply install the [NuGet package](https://www.nuget.org/packages/Mindee) using `dotnet add`:

```sh
dotnet add package Mindee --version 4.4
```

{% endtab %}
{% endtabs %}

Don't see support for your favorite language or framework? [Make a feature request!](https://feedback.mindee.com/?b=682f69c9e2404756e7e68d1c)

## Send a File and Poll

Use the client library to send the file to your Crop Model and return the result.

{% tabs %}
{% tab title="Python" %}
Requires Python ≥ 3.9. Python ≥ 3.11 is recommended.\
Requires the [Mindee Python SDK](https://pypi.org/project/mindee/) version **4.35.1** or greater.

{% code lineNumbers="true" %}

```python
from mindee import PathInput
from mindee.v2 import (
    Client,
    CropParameters,
    CropResponse,
)

input_path = "/path/to/the/file.ext"
api_key = "MY_API_KEY"
model_id = "MY_MODEL_ID"

# Init a new client
mindee_client = Client(api_key)

# Set Crop parameters
model_params = CropParameters(
    # ID of the model, required.
    model_id=model_id,
)

# Load a file from disk
input_source = PathInput(input_path)

# Send for processing using polling
response = mindee_client.enqueue_and_get_result(
    CropResponse,
    input_source,
    model_params,
)

# Print a brief summary of the parsed data
print(response.inference)

# Access the crop result
crops: list = response.inference.result.crops
```

{% endcode %}

Also take a look at the [Crop Result](https://docs.mindee.com/crop-models/sdk-integration/crop-result) documentation.
{% endtab %}

{% tab title="Node.js" %}
Requires Node.js ≥ 20.1. Node.js ≥ 22 is recommended.\
Requires the [Mindee Node.js SDK](https://www.npmjs.com/package/mindee/) version **5.5.0** or greater.

{% code lineNumbers="true" %}

```javascript
import * as mindee from "mindee";
// If you're on CommonJS:
// const mindee = require("mindee");

const apiKey = "MY_API_KEY";
const filePath = "/path/to/the/file.ext";
const modelId = "MY_MODEL_ID";

// Init a new client
const mindeeClient = new mindee.Client(
  { apiKey: apiKey }
);

// Set Crop parameters
const modelParams = {
  modelId: modelId,
};

// Load a file from disk
const inputSource = new mindee.PathInput({ inputPath: filePath });

// Send for processing
const response = await mindeeClient.enqueueAndGetResult(
  mindee.product.Crop,
  inputSource,
  modelParams,
);

// print a string summary
console.log(response.inference.toString());

// Access the result crops
const crops = response.inference.result.crops;
```

{% endcode %}

Also take a look at the [Crop Result](https://docs.mindee.com/crop-models/sdk-integration/crop-result) documentation.
{% endtab %}

{% tab title="PHP" %}
Requires PHP ≥ 8.1. PHP ≥ 8.3 is recommended.\
Requires the [Mindee PHP SDK](https://packagist.org/packages/mindee/mindee) version **3.0.0** or greater.

{% code lineNumbers="true" %}

```php
<?php

use Mindee\Input\PathInput;
use Mindee\V2\Client;
use Mindee\V2\Product\Crop\Params\CropParameters;
use Mindee\V2\Product\Crop\CropResponse;

$apiKey = "MY_API_KEY";
$modelId = "MY_MODEL_ID";
$filePath = "/path/to/the/file.ext";

// Init a new client
$mindeeClient = new Client($apiKey);

// Set Crop parameters
$modelParams = new CropParameters(
    // ID of the model, required.
    $modelId,
);

// Load a file from disk
$inputSource = new PathInput($filePath);

// Send for processing using polling
$response = $mindeeClient->enqueueAndGetResult(
    CropResponse::class,
    $inputSource,
    $modelParams
);

// Print a summary of the response
echo strval($response->inference);

// Access the crop results
$crops = $response->inference->result->crops;
```

{% endcode %}

Also take a look at the [Crop Result](https://docs.mindee.com/crop-models/sdk-integration/crop-result) documentation.
{% endtab %}

{% tab title="Ruby" %}
Requires Ruby ≥ 3.2.\
Requires the [Mindee Ruby SDK](https://rubygems.org/gems/mindee) version **5.2.1** or greater.

{% code lineNumbers="true" %}

```ruby
require 'mindee'
require 'mindee/v2/product'

input_path = '/path/to/the/file.ext'
api_key = 'MY_API_KEY'
model_id = 'MY_MODEL_ID'

# Init a new client
mindee_client = Mindee::V2::Client.new(api_key: api_key)

# Set Crop parameters
model_params = {
    # ID of the model, required.
    model_id: model_id,
}

# Load a file from disk
input_source = Mindee::Input::Source::PathInputSource.new(input_path)

# Send for processing
response = mindee_client.enqueue_and_get_result(
    Mindee::V2::Product::Crop::Crop,
    input_source,
    model_params
)

# Access the result crops
puts response.inference.result.crops
```

{% endcode %}

Also take a look at the [Crop Result](https://docs.mindee.com/crop-models/sdk-integration/crop-result) documentation.
{% endtab %}

{% tab title="Java" %}
Requires Java ≥ 11. Java ≥ 17 is recommended.\
Requires the [Mindee Java SDK](https://central.sonatype.com/artifact/com.mindee.sdk/mindee-api-java) version **5.1.0** or greater.

{% code lineNumbers="true" %}

```java
import com.mindee.input.LocalInputSource;
import com.mindee.v2.MindeeClient;
import com.mindee.v2.product.crop.CropResponse;
import com.mindee.v2.product.crop.params.CropParameters;
import java.io.IOException;

public class SimpleMindeeClientV2 {

  public static void main(String[] args)
      throws IOException, InterruptedException
  {
    String apiKey = "MY_API_KEY";
    String modelId = "MY_MODEL_ID";
    String filePath = "/path/to/the/file.ext";

    // Init a new client
    var mindeeClient = new MindeeClient(apiKey);

    // Set Crop parameters
    var modelParams = CropParameters
        // ID of the model, required.
        .builder(modelId)
        .build();

    // Load a file from disk
    var inputSource = new LocalInputSource(filePath);

    // Send for processing using polling
    CropResponse response = mindeeClient.enqueueAndGetResult(
        CropResponse.class,
        inputSource,
        modelParams
    );

    // Print a summary of the response
    System.out.println(response.getInference().toString());

    // Access the crop results
    var crops = response.getInference().getResult().getCrops();
  }
}
```

{% endcode %}

Also take a look at the [Crop Result](https://docs.mindee.com/crop-models/sdk-integration/crop-result) documentation.
{% endtab %}

{% tab title=".NET" %}
.NET ≥ 8.0 is recommended.\
Requires the [Mindee .NET SDK](https://www.nuget.org/packages/Mindee) version **4.3.0** or greater.

{% code lineNumbers="true" %}

```csharp
using Mindee;
using Mindee.Input;
using Mindee.V2;
using Mindee.V2.Product.Crop;
using Mindee.V2.Product.Crop.Params;

string filePath = "/path/to/the/file.ext";
string apiKey = "MY_API_KEY";
string modelId = "MY_MODEL_ID";

// Construct a new client
Client mindeeClient = new Client(apiKey);

// Set Crop parameters
var modelParams = new CropParameters(
    modelId: modelId
);

// Load a file from disk
var inputSource = new LocalInputSource(filePath);

// Upload the file
var response = await mindeeClient.EnqueueAndGetResultAsync<CropResponse>(
    inputSource, modelParams);

// Print a summary of the response
System.Console.WriteLine(response.Inference.ToString());

// Access the crop results
var crops = response.Inference.Result.Crops;
```

{% endcode %}

Also take a look at the [Crop Result](https://docs.mindee.com/crop-models/sdk-integration/crop-result) documentation.
{% endtab %}
{% endtabs %}


# Crop Result

Reference documentation on processing a Crop result using the Mindee SDKs.

You'll need to have a response object as described in the [Response Processing](/integrations/client-libraries-sdk/process-the-response) section.

## Accessing Crop Items

A `CropItem` describes the location of a single detected object.

In this context, an *object* can be any element to find on a page.

Most users will look for a document type (like a receipt or an ID), but it could really be anything (a photo, a logo, etc).&#x20;

### `CropItem` Attributes

#### Object Type

The category assigned to the object. It is always filled.

It is returned as a string and is identical to the value entered on the Mindee platform — case, spaces, and punctuation included.

Why "Object Type" instead of "Document Type", like for Split and Classification? Because the fundamental technology is different, Crop uses Object Detection algorithms, whereas Split and Classification use variations of classification technology.

#### Location

The location of the object in the document. It contains the following properties:

* Polygon: Coordinates of the object.
* Page: 0-based index of the page the coordinates were found on.

#### Extraction Response

Optional extraction response associated with the split. This is only filled if extraction chaining is activated for the model.

### Iterate Over Crop Items

You'll usually want to iterate over all crop items, since the number of items is dependent on the document. Remember that a document can have multiple pages, and each of its pages can have multiple crop items.

{% tabs %}
{% tab title="Python" %}

```python
from mindee import CropResponse

def handle_response(response: CropResponse) -> None:
    crops = response.inference.result.crops

    for crop in crops:
        # Object type identified for this crop
        object_type = crop.object_type
        print(f"Detected type: {object_type}")

        # Location of the crop
        location = crop.location

        # 0-based page index
        page = location.page

        # Polygon object, which includes some useful methods
        polygon = location.polygon
        center = polygon.centroid

        print(f"On page {page}, with center at {center}")

        # Optional extraction response, present if extraction chaining was requested
        extraction_response = crop.extraction_response

        if extraction_response is not None:
            # Access extracted fields from the crop's inference result
            fields = extraction_response.inference.result.fields
            print(f"Extraction fields: {fields}")
```

{% endtab %}

{% tab title="Node.js" %}

```javascript
function handleResponse(response): void {
  const crops = response.inference.result.crops;

  for (const crop of crops) {
    // Object type identified for this crop
    const objectType = crop.objectType;
    console.log(`Detected type: ${objectType}`);

    // Location of the crop
    const location = crop.location;

    // 0-based page index
    const page = location.page;

    // Polygon object, which includes some useful methods
    const polygon = location.polygon;
    const center = polygon.getCentroid();

    console.log(`On page ${page}, with center at ${center}`);

    // Optional extraction response, present if extraction chaining was requested
    const extractionResponse = crop.extractionResponse;

    if (extractionResponse !== undefined) {
      // Access extracted fields from the crop's inference result
      const fields = extractionResponse.inference.result.fields;
      console.log(`Extraction fields: ${fields}`);
    }
  }
}
```

{% endtab %}

{% tab title="PHP" %}

```php
use Mindee\V2\Product\Split\CropResponse;

public function handleResponse(CropResponse $response)
{
    $crops = $response->inference->result->crops;
}
```

{% endtab %}

{% tab title="Ruby" %}

```ruby
def handle_response(response)
  crops = response.inference.result.crops

  crops.each do |crop|
    # Object type identified for this crop
    object_type = crop.object_type
    puts "Detected type: #{object_type}"

    # Location of the crop
    location = crop.location

    # 0-based page index
    page = location.page

    # Polygon object, which includes some useful methods
    polygon = location.polygon
    center = polygon.centroid

    puts "On page #{page}, with center at #{center}"

    # Optional extraction response, present if extraction chaining was requested
    extraction_response = crop.extraction_response

    unless extraction_response.nil?
      # Access extracted fields from the split's inference result
      fields = extraction_response.inference.result.fields
      puts "Extraction fields: #{fields}"
    end
  end
end
```

{% endtab %}

{% tab title="Java" %}

```java
import com.mindee.v2.product.crop.CropResponse;

public void handleResponse(CropResponse response) {
  var crops = response.getInference().getResult().getCrops();
  
  for (var crop : crops) {
    // Object type identified for this crop
    String objectType = crop.getObjectType();
    System.out.println("Detected type: " + objectType);

    // location of the crop
    var location = crop.getLocation();

    // 0-based page index
    var page = location.getPage();

    // polygon object, which includes some useful methods
    var polygon = location.getPolygon();
    var center = polygon.getCentroid();

    System.out.println("On page " + page + ", with center at " + center);

    // Optional extraction response, present if extraction chaining was requested
    var extractionResponse = crop.getExtractionResponse();

    if (extractionResponse != null) {
      // Access extracted fields from the split's inference result
      var fields = extractionResponse.getInference().getResult().getFields();
      System.out.println("Extraction fields: " + fields.toString());
    }
  }
}
```

{% endtab %}

{% tab title=".NET" %}

```csharp
using Mindee.V2.Product.Crop;

public void HandleResponse(CropResponse response)
{
    var crops = response.Inference.Result.Crops;

    foreach (var crop in crops)
    {
        // Object type identified for this crop
        string objectType = crop.ObjectType;
        Console.WriteLine($"Detected type: {objectType}");

        // Location of the crop
        var location = crop.Location;

        // 0-based page index
        var page = location.Page;

        // Polygon object, which includes some useful methods
        var polygon = location.Polygon;
        var center = polygon.GetCentroid();

        Console.WriteLine($"On page {page}, with center at {center}");

        // Optional extraction response, present if extraction chaining was requested
        var extractionResponse = crop.ExtractionResponse;

        if (extractionResponse != null)
        {
            // Access extracted fields from the crop's inference result
            var fields = extractionResponse.Inference.Result.Fields;
            Console.WriteLine($"Extraction fields: {fields}");
        }
    }
}
```

{% endtab %}
{% endtabs %}

Chained Extraction results will be an `ExtractionResponse` object. This object is identical to the response of an Extraction Model.

Refer to the [Extraction Result](/extraction-models/sdk-integration/extraction-result) section for details on processing the Extraction.


# No-Code Integration

No-Code and Low-Code integration support for Crop models.

## Officially Supported Integrations

We're hard at work providing official integrations for Crop models on platforms that allow it.

Currently we have support for the following platforms:

* **n8n**: official support [coming soon](https://github.com/n8n-io/n8n/pull/18986)
* **Zapier**: officially supported
* **Make**: officially supported

Don't see support for your favorite platform? [Make a feature request!](https://feedback.mindee.com/?b=682f69c9e2404756e7e68d1c)

{% hint style="info" %}
**Usage is very similar to** [**Extraction models**](/extraction-models/no-code-integration)**.**

Complete documentation in progress...
{% endhint %}


# Classification Model Overview

Automatically attribute a class to a given document, from a list of classes you define yourself.

## Use Cases

Process a single document sent in a single file. The classifier looks at all pages of the document in order to identify its type.

Some common examples:

* You have multiple types of files in your workflow input, with different business rules
* You want to identify the region or language of documents

A file sent to the Classification Model may have any number of pages, [within limits](/integrations/technical-limitations#file-limits).

{% hint style="info" icon="lightbulb" %}
If there is a high possibility of having multiple documents within the same file, use:

* [Split Model Overview](/split-models/split) ⇒ multiple documents in the same file
* [Crop Model Overview](/crop-models/crop) ⇒ multiple documents on the same page
  {% endhint %}

## Create a Classification Model

Classification models are always custom, there are no templates available in the Catalog.\
This keeps Classification models flexible for different documents and workflows.

Each Classification model gets its own unique model ID when you create it.

1. To create a Classification model, click on **Models**, and then click on **Create your document AI model**.
2. Scroll to the **Document Utilities** section, click on **Classify.**
3. A pop-up will appear, allowing you to enter the classes you want. Most of the time, you'll use one possible document type per class.\
   \
   For example, if the files you are processing contain invoices, receipts, and driving licenses, set the classes as:\
   `INVOICES`, `IDENTITY DOCUMENTS`, `CONTRACTS`.\
   \
   Classes are returned exactly as defined: including upper or lower case, and any spaces.

{% hint style="info" icon="lightbulb" %}
Add the class `OTHER` if you need the model to identify documents that are not one of the explicitly defined classes.
{% endhint %}

<figure><img src="/files/EvjxXfTENfXZfbXxwMqT" alt="" width="375"><figcaption></figcaption></figure>

4. Once ready, click on **Create Utility** to create your Classification Model.\
   This step will also generate the model's unique ID.
5. You can now use the **Live Test** tab to process documents, and the **Utility Configuration** to update your classes.<br>

Your utility is now available in your **Models** tab.

<figure><img src="/files/46NeM7yr2tM9tbYTVqok" alt=""><figcaption></figcaption></figure>

Here is a step-by-step tutorial that shows you how to properly create a Classify Utility :

{% @supademo/embed url="<https://app.supademo.com/demo/cmls1rubd1jal11896wj0sg8q>" demoId="cmls1rubd1jal11896wj0sg8q" %}

## Technical Considerations

Models requires at least two classes defined.

Class names may be in most languages and writing systems, but cannot exceed 128 characters.

Having more than a few dozen classes will yield unexpected results, and is not recommended. If you need classes based on vendor/customer names, product codes, phone numbers, etc, you should use a text field in an [Extraction model](/extraction-models/extraction-models-overview) instead.

{% hint style="info" %}
Class names will be returned **exactly** as defined on the platform in the return, spaces and capitalization included.

If the class names are changed on the platform, the change in the API return will be **immediate** for all new files sent.
{% endhint %}

## Integration

Once your Classification model is created and tested, integration documentation is provided in the "Documentation" page, or here: [Classification Quick Start](/classification-models/sdk-integration/classification-quick-start).


# Extraction Model Chaining

Extract data from detected classes.

Use Classification to automatically extract the correct data from files.

Note: Classification Models always return a single class regardless of the number of pages or if there are multiple documents on the same page.

## Extraction Set Up

### At Model Creation

When creating your Classification Model, you'll be adding document classes in the creation window.

For each document class, you can link one of your [Extraction Models](/extraction-models/extraction-models-overview) for chaining. The Extraction Model must exist prior to the Classification Model creation.

Use the search field to filter available Extraction Models.

### After Model Creation

This works exactly like when creating at model creation.

Simply go to your Split model's "Utility Configuration" page and adjust as needed.

You can add new classes, remove classes, and change Extraction Models.

### Selectively Extracting

If a detected class has no linked Extraction Model, no extraction runs for that class.

This allows selectively extracting some files while ignoring others.

Let's say you receive mixed file types from your users, typically but not limited to: plane tickets, travel receipts, driver licenses, and passports.

You need only passports. In your Classification configuration, add a `passport` class and an `other`  class, and only link an extraction model to the `passport` class.

All documents will get classified, but only those linked to an Extraction Model will have extraction results.

{% hint style="info" icon="lightbulb" %}
Add an `other` class in addition to classes of interest to you, unless all potential document types are known in advance.
{% endhint %}

### Token Usage With Chaining

You only consume tokens once, there is no double charge when chaining to an Extraction model.

The possible credit consumption scenarios are as follows:

* extraction model not chained ⇒ token usage of initial model (Split, Crop, Classification)
* extraction model chained ⇒ token usage of the configured extraction model, including any activated [optional features](/extraction-models/optional-features).

In other words, the initial model is provided at no additional cost to you if it triggers additional processing using an Extraction model.

All token consumption is per page as is standard for our models.

## Access Extraction Results

When an Extraction Model is linked, the detected class contains an Extraction Response object.

That object is the same as the response returned by a standalone Extraction request.

Check [Classification Result](/classification-models/sdk-integration/classification-result) for details on accessing crop items and their extraction responses.


# SDK Integration

Integrate a Classification model using the Mindee SDKs.

Use the SDKs to send documents to a Classification model and process classification results.

This section helps you choose the right starting point and move through the full integration flow.

### Choose your path

Start with the page that matches your next step:

* [Classification Quick Start](/classification-models/sdk-integration/classification-quick-start) ⇒ install a client library, send a file, and get your first result.
* [Classification Result](/classification-models/sdk-integration/classification-result) ⇒ access the predicted class and optional chained extraction result.
* [Classification Model Overview](/classification-models/classification) ⇒ understand how Classification models assign a class to a document.

### Typical SDK workflow

{% stepper %}
{% step %}

#### Create Your Model

[Create your Classification Model](/classification-models/classification#create-a-classification-model) on the Mindee platform.

Upload some samples to the [Live Test](/models/live-test) to validate the model.
{% endstep %}

{% step %}

### Install and authenticate

Install the client library for your language.

Prepare your API key and initialize the client.
{% endstep %}

{% step %}

#### Send a document

Set the Classification model's ID and send a file or URL for processing.

Start with polling unless you already use webhooks.
{% endstep %}

{% step %}

#### Process the classification result

Read the predicted class from the response.

Use the returned document type to route the document in your workflow.
{% endstep %}
{% endstepper %}

### Before you start

Have these ready:

* your API key
* your Classification Model's unique ID
* a sample document for testing
* a language choice for the SDK

### What is specific to Classification models

Classification responses describe the class predicted for the whole document.

The classifier looks at all pages in the file before assigning a document type.

Document types are returned exactly as configured on the platform.

If extraction chaining is enabled, the classification result can also include an extraction response.

### Shared SDK building blocks

Integration builds on the same client library concepts used across all model types.

* [Client Libraries / SDKs](/integrations/client-libraries-sdk)
* [Client Configuration](/integrations/client-libraries-sdk/configure-the-client)
* [Send a File or URL](/integrations/client-libraries-sdk/send-a-file-or-url)
* [Response Processing](/integrations/client-libraries-sdk/process-the-response)
* [Webhook Results](/integrations/webhooks)
* [Error Handling](/integrations/problem-database)


# Classification Quick Start

Integrate a Classification Model.

## Install the Client Library

Install the Mindee SDK for your language or framework of choice

{% tabs %}
{% tab title="Python" %}
Requires Python ≥ 3.9. Python ≥ 3.11 is recommended.

Simply install the [PyPi package](https://pypi.org/project/mindee/) using `pip`:

```sh
pip install -U mindee~=5.2
```

{% endtab %}

{% tab title="Node.js" %}
Requires Node.js ≥ 20.1. Node.js ≥ 22 is recommended.

Simply install the [NPM package](https://www.npmjs.com/package/mindee):

```sh
npm install mindee@^5.5.0
```

{% endtab %}

{% tab title="PHP" %}
Requires PHP ≥ 8.1. PHP ≥ 8.3 is recommended.

Simply install the [Packagist package](https://packagist.org/packages/mindee/mindee) using [composer](https://getcomposer.org/):

```sh
php composer.phar require "mindee/mindee:>=3.0"
```

{% endtab %}

{% tab title="Ruby" %}
Requires Ruby ≥ 3.2.

Simply install the [gem](https://rubygems.org/gems/mindee) using:

```shell
gem install mindee -v '~> 5.2'
```

{% endtab %}

{% tab title="Java" %}
Requires Java ≥ 11. Java ≥ 17 is recommended.

Group ID: `com.mindee.sdk`\
Artifact ID: `mindee-api-java`\
Version: `5.2.0` or greater

There are various installation methods, Maven, Gradle, etc:

[Installation Details](https://central.sonatype.com/artifact/com.mindee.sdk/mindee-api-java)
{% endtab %}

{% tab title=".NET" %}
.NET ≥ 8.0 is recommended.

Simply install the [NuGet package](https://www.nuget.org/packages/Mindee) using `dotnet add`:

```sh
dotnet add package Mindee --version 4.4
```

{% endtab %}
{% endtabs %}

Don't see support for your favorite language or framework? [Make a feature request!](https://feedback.mindee.com/?b=682f69c9e2404756e7e68d1c)

## Send a File and Poll

Use the client library to send the file to your Classification Model and return the result.

{% tabs %}
{% tab title="Python" %}
Requires Python ≥ 3.9. Python ≥ 3.11 is recommended.\
Requires the [Mindee Python SDK](https://pypi.org/project/mindee/) version &#x35;**.2.0** or greater.

{% code lineNumbers="true" %}

```python
from mindee import PathInput
from mindee.v2 import (
    Client,
    ClassificationParameters,
    ClassificationResponse,
)

input_path = "/path/to/the/file.ext"
api_key = "MY_API_KEY"
model_id = "MY_MODEL_ID"

# Init a new client
mindee_client = Client(api_key)

# Set Classification parameters
model_params = ClassificationParameters(
    # ID of the model, required.
    model_id=model_id,
)

# Load a file from disk
input_source = PathInput(input_path)

# Send for processing using polling
response = mindee_client.enqueue_and_get_result(
    ClassificationResponse,
    input_source,
    model_params,
)

# Print a brief summary of the parsed data
print(response.inference)

# Access the classification result
classification: str = response.inference.result.classification.document_type
```

{% endcode %}

Also take a look at the [Classification Result](https://docs.mindee.com/classification-models/sdk-integration/classification-result) documentation.
{% endtab %}

{% tab title="Node.js" %}
Requires Node.js ≥ 20.1. Node.js ≥ 22 is recommended.\
Requires the [Mindee Node.js SDK](https://www.npmjs.com/package/mindee/) version **5.5.0** or greater.

{% code lineNumbers="true" %}

```javascript
import * as mindee from "mindee";
// If you're on CommonJS:
// const mindee = require("mindee");

const apiKey = "MY_API_KEY";
const filePath = "/path/to/the/file.ext";
const modelId = "MY_MODEL_ID";

// Init a new client
const mindeeClient = new mindee.Client(
  { apiKey: apiKey }
);

// Set Classification parameters
const modelParams = {
  modelId: modelId,
};

// Load a file from disk
const inputSource = new mindee.PathInput({ inputPath: filePath });

// Send for processing
const response = await mindeeClient.enqueueAndGetResult(
  mindee.product.Classification,
  inputSource,
  modelParams,
);

// print a string summary
console.log(response.inference.toString());

// Access the classification result
const classification = response.inference.result.classification;
```

{% endcode %}

Also take a look at the [Classification Result](https://docs.mindee.com/classification-models/sdk-integration/classification-result) documentation.
{% endtab %}

{% tab title="PHP" %}
Requires PHP ≥ 8.1. PHP ≥ 8.3 is recommended.\
Requires the [Mindee PHP SDK](https://packagist.org/packages/mindee/mindee) version **3.0.0** or greater.

{% code lineNumbers="true" %}

```php
<?php

use Mindee\Input\PathInput;
use Mindee\V2\Client;
use Mindee\V2\Product\Classification\Params\ClassificationParameters;
use Mindee\V2\Product\Classification\ClassificationResponse;

$apiKey = "MY_API_KEY";
$modelId = "MY_MODEL_ID";
$filePath = "/path/to/the/file.ext";

// Init a new client
$mindeeClient = new Client($apiKey);

// Set Classification parameters
$modelParams = new ClassificationParameters(
    // ID of the model, required.
    $modelId,
);

// Load a file from disk
$inputSource = new PathInput($filePath);

// Send for processing using polling
$response = $mindeeClient->enqueueAndGetResult(
    ClassificationResponse::class,
    $inputSource,
    $modelParams
);

// Print a summary of the response
echo strval($response->inference);

// Access the classification results
$classification = $response->inference->result->classification;
```

{% endcode %}

Also take a look at the [Classification Result](https://docs.mindee.com/classification-models/sdk-integration/classification-result) documentation.
{% endtab %}

{% tab title="Ruby" %}
Requires Ruby ≥ 3.2.\
Requires the [Mindee Ruby SDK](https://rubygems.org/gems/mindee) version **5.2.1** or greater.

{% code lineNumbers="true" %}

```ruby
require 'mindee'
require 'mindee/v2/product'

input_path = '/path/to/the/file.ext'
api_key = 'MY_API_KEY'
model_id = 'MY_MODEL_ID'

# Init a new client
mindee_client = Mindee::V2::Client.new(api_key: api_key)

# Set Classification parameters
model_params = {
    # ID of the model, required.
    model_id: model_id,
}

# Load a file from disk
input_source = Mindee::Input::Source::PathInputSource.new(input_path)

# Send for processing
response = mindee_client.enqueue_and_get_result(
    Mindee::V2::Product::Classification::Classification,
    input_source,
    model_params
)

# Access the classification result
puts response.inference.result.classification
```

{% endcode %}

Also take a look at the [Classification Result](https://docs.mindee.com/classification-models/sdk-integration/classification-result) documentation.
{% endtab %}

{% tab title="Java" %}
Requires Java ≥ 11. Java ≥ 17 is recommended.\
Requires the [Mindee Java SDK](https://central.sonatype.com/artifact/com.mindee.sdk/mindee-api-java) version **5.1.0** or greater.

{% code lineNumbers="true" %}

```java
import com.mindee.input.LocalInputSource;
import com.mindee.v2.MindeeClient;
import com.mindee.v2.product.classification.ClassificationResponse;
import com.mindee.v2.product.classification.params.ClassificationParameters;
import java.io.IOException;

public class SimpleMindeeClientV2 {

  public static void main(String[] args)
      throws IOException, InterruptedException
  {
    String apiKey = "MY_API_KEY";
    String modelId = "MY_MODEL_ID";
    String filePath = "/path/to/the/file.ext";

    // Init a new client
    var mindeeClient = new MindeeClient(apiKey);

    // Set Classification parameters
    var modelParams = ClassificationParameters
        // ID of the model, required.
        .builder(modelId)
        .build();

    // Load a file from disk
    var inputSource = new LocalInputSource(filePath);

    // Send for processing using polling
    ClassificationResponse response = mindeeClient.enqueueAndGetResult(
        ClassificationResponse.class,
        inputSource,
        modelParams
    );

    // Print a summary of the response
    System.out.println(response.getInference().toString());

    // Access the classification result
    var result = response.getInference().getResult();
    String documentType = result.getClassification().getDocumentType();
  }
}
```

{% endcode %}

Also take a look at the [Classification Result](https://docs.mindee.com/classification-models/sdk-integration/classification-result) documentation.
{% endtab %}

{% tab title=".NET" %}
.NET ≥ 8.0 is recommended.\
Requires the [Mindee .NET SDK](https://www.nuget.org/packages/Mindee) version **4.3.0** or greater.

{% code lineNumbers="true" %}

```csharp
using Mindee;
using Mindee.Input;
using Mindee.V2;
using Mindee.V2.Product.Classification;
using Mindee.V2.Product.Classification.Params;

string filePath = "/path/to/the/file.ext";
string apiKey = "MY_API_KEY";
string modelId = "MY_MODEL_ID";

// Construct a new client
Client mindeeClient = new Client(apiKey);

// Set inference parameters
var modelParams = new ClassificationParameters(
    modelId: modelId
);

// Load a file from disk
var inputSource = new LocalInputSource(filePath);

// Upload the file
var response = await mindeeClient.EnqueueAndGetResultAsync<ClassificationResponse>(
    inputSource, modelParams);

// Print a summary of the response
System.Console.WriteLine(response.Inference.ToString());

// Access the classification results
var classification = response.Inference.Result.Classification;
```

{% endcode %}

Also take a look at the [Classification Result](https://docs.mindee.com/classification-models/sdk-integration/classification-result) documentation.
{% endtab %}
{% endtabs %}


# Classification Result

Reference documentation on processing a Classification result using the Mindee SDKs.

You'll need to have a response object as described in the [Response Processing](/integrations/client-libraries-sdk/process-the-response) section.

## Accessing the Classification

A `Classifier` describes the detected class of the entire document.

### `Classifier` Attributes

#### Document Type

The document category assigned to the sub-document. It is always filled.

It is returned as a string and is identical to the value entered on the Mindee platform — case, spaces, and punctuation included.

#### Extraction Response

Optional extraction response associated with the split. This is only filled if extraction chaining is activated for the model.

{% tabs %}
{% tab title="Python" %}

```python
from mindee import ClassificationResponse

def handle_response(response: ClassificationResponse) -> None:
    classification = response.inference.result.classification

    # Document type identified for this file
    document_type: str = classification.document_type
    print(f"Detected type: {document_type}")

    # Optional extraction response, present if extraction chaining was requested
    extraction_response = classification.extraction_response

    if extraction_response is not None:
        # Access extracted fields from the inference result
        fields = extraction_response.inference.result.fields
        print(f"Extraction fields: {fields}")
```

{% endtab %}

{% tab title="Node.js" %}

```javascript
function handleResponse(response) {
  const classification = response.getInference().getResult().getClassification();

  // Document type identified for this file
  const documentType = classification.getDocumentType();
  console.log(`Detected type: ${documentType}`);

  // Optional extraction response, present if extraction chaining was requested
  const extractionResponse = classification.getExtractionResponse();

  if (extractionResponse != null) {
    // Access extracted fields from the split's inference result
    const fields = extractionResponse.getInference().getResult().getFields();
    console.log(`Extraction fields: ${fields.toString()}`);
  }
}
```

{% endtab %}

{% tab title="PHP" %}

```php
use Mindee\V2\Product\Classification\ClassificationResponse;

public function handleResponse(SplitResponse $response)
{
    $classification = $response->inference->result->classification;
}
```

{% endtab %}

{% tab title="Ruby" %}

```ruby
def handle_response(response)
  classification = response.inference.result.classification

  # Document type identified for this file
  document_type = classification.document_type
  puts "Detected type: #{document_type}"

  # Optional extraction response, present if extraction chaining was requested
  extraction_response = classification.extraction_response

  unless extraction_response.nil?
    # Access extracted fields from the split's inference result
    fields = extraction_response.inference.result.fields
    puts "Extraction fields: #{fields}"
  end
end
```

{% endtab %}

{% tab title="Java" %}

```java
import com.mindee.v2.product.classification.ClassificationResponse;

public void handleResponse(ClassificationResponse response) {
  var classification = response.getInference().getResult().getClassification();
  
  // Document type identified for this file
  String documentType = classification.getDocumentType();
  System.out.println("Detected type: " + documentType);

  // Optional extraction response, present if extraction chaining was requested
  var extractionResponse = classification.getExtractionResponse();

  if (extractionResponse != null) {
    // Access extracted fields from the split's inference result
    var fields = extractionResponse.getInference().getResult().getFields();
    System.out.println("Extraction fields: " + fields.toString());
  }
}
```

{% endtab %}

{% tab title=".NET" %}

```csharp
using Mindee.V2.Product.Classification;

public void HandleResponse(ClassificationResponse response)
{
    var classification = response.Inference.Result.Classification;

    // Document type identified for this file
    string documentType = classification.DocumentType;
    Console.WriteLine($"Detected type: {documentType}");

    // Optional extraction response, present if extraction chaining was requested
    var extractionResponse = classification.ExtractionResponse;

    if (extractionResponse != null)
    {
        // Access extracted fields from the classification's inference result
        var fields = extractionResponse.Inference.Result.Fields;
        Console.WriteLine($"Extraction fields: {fields}");
    }
}
```

{% endtab %}
{% endtabs %}

Chained Extraction results will be an `ExtractionResponse` object. This object is identical to the response of an Extraction Model.

Refer to the [Extraction Result](/extraction-models/sdk-integration/extraction-result) section for details on processing the Extraction.


# No-Code Integration

No-Code and Low-Code integration support for Classification models.

## Officially Supported Integrations

We're hard at work providing official integrations for Classification models on platforms that allow it.

Currently we have support for the following platforms:

* **n8n**: official support [coming soon](https://github.com/n8n-io/n8n/pull/18986)
* **Zapier**: officially supported
* **Make**: officially supported

Don't see support for your favorite platform? [Make a feature request!](https://feedback.mindee.com/?b=682f69c9e2404756e7e68d1c)

{% hint style="info" %}
**Usage is very similar to** [**Extraction models**](/extraction-models/no-code-integration)**.**

Complete documentation in progress...
{% endhint %}


# Raw Text Model Overview

Automatically extract the raw text from each page of a document using OCR.

## Use Cases

When the entire text of a document needs to be extracted, along with position information for each word.

We use the term "OCR" literally, as the acronym for "Optical Character Recognition".

Meaning the exact text is returned as raw data, not parsed into structured data as fields.

{% hint style="info" %}
**Most users looking for a generic "Mindee OCR" or to "OCR a document" are likely looking for Extraction: c**onsult the [Extraction model documentation](/extraction-models/extraction-models-overview).
{% endhint %}

A file sent to the Raw Text OCR Model may have any number of pages, [within limits](/integrations/technical-limitations#file-limits).

### Difference Between Raw Text *Model* and Raw Text *Option*

The [Extraction model](/extraction-models/extraction-models-overview) has the [Raw Text optional feature](/extraction-models/optional-features/raw-text-full-ocr) that can be used in similar uses cases, depending on your needs.

**Raw Text option:** returns the entire text for each page as a single string.

**Raw Text model:** in addition to the single string per page, also returns each word on the page. Each word has position information and text content.

### Language Support

Almost all languages are fully supported, since only the writing system is involved in detection.

In other words, as long as the system can recognize the glyphs (letters in an alphabet), the text will be extracted.

It's much easier and shorter to list what is **not** supported:

* Ancient languages with no equivalent modern writing system such as cuneiform, Egyptian hieroglyphs, ancient Maya, etc.
* Modern languages with uncommon writing systems such as Blackfoot, Cherokee, Inuktitut, etc.

## Create a Raw Text OCR Model

Raw OCR models are always custom, there are no templates available in the Catalog.

Each OCR model gets its own unique model ID when you create it.

1. To create a OCR model, click on **Models**, and then click on **Create your document AI model**.
2. Scroll to the **Document Utilities** section, click on **OCR.**\
   This step will also generate the model's unique ID.
3. You can now use the **Live Test** tab to process documents.<br>

Your OCR model is now available in your **Models** tab:

<figure><img src="/files/YThWXqYUf397uDSxX7wv" alt=""><figcaption></figcaption></figure>

Here is a step-by-step tutorial that shows you how to properly create an OCR utility:

{% @supademo/embed url="<https://app.supademo.com/demo/cmlrrputf0pnh1189ay3tsdkt>" demoId="cmlrrputf0pnh1189ay3tsdkt" %}

## Integration

Once your OCR model is created and tested, integration documentation is provided in the "Documentation" page, or here: [OCR Quick Start](/raw-text-ocr-models/sdk-integration/ocr-quick-start)


# SDK Integration

Integrate an OCR model using the Mindee SDKs.


# OCR Quick Start

Integrate an OCR Model.

## Install the Client Library

Install the Mindee Client Library for your language or framework of choice

{% tabs %}
{% tab title="Python" %}
Requires Python ≥ 3.9. Python ≥ 3.11 is recommended.

Simply install the [PyPi package](https://pypi.org/project/mindee/) using `pip`:

```sh
pip install -U mindee~=5.2
```

{% endtab %}

{% tab title="Node.js" %}
Requires Node.js ≥ 20.1. Node.js ≥ 22 is recommended.

Simply install the [NPM package](https://www.npmjs.com/package/mindee):

```sh
npm install mindee@^5.5.0
```

{% endtab %}

{% tab title="PHP" %}
Requires PHP ≥ 8.1. PHP ≥ 8.3 is recommended.

Simply install the [Packagist package](https://packagist.org/packages/mindee/mindee) using [composer](https://getcomposer.org/):

```sh
php composer.phar require "mindee/mindee:>=3.0"
```

{% endtab %}

{% tab title="Ruby" %}
Requires Ruby ≥ 3.2.

Simply install the [gem](https://rubygems.org/gems/mindee) using:

```shell
gem install mindee -v '~> 5.2'
```

{% endtab %}

{% tab title="Java" %}
Requires Java ≥ 11. Java ≥ 17 is recommended.

Group ID: `com.mindee.sdk`\
Artifact ID: `mindee-api-java`\
Version: `5.2.0` or greater

There are various installation methods, Maven, Gradle, etc:

[Installation Details](https://central.sonatype.com/artifact/com.mindee.sdk/mindee-api-java)
{% endtab %}

{% tab title=".NET" %}
.NET ≥ 8.0 is recommended.

Simply install the [NuGet package](https://www.nuget.org/packages/Mindee) using `dotnet add`:

```sh
dotnet add package Mindee --version 4.4
```

{% endtab %}
{% endtabs %}

Don't see support for your favorite language or framework? [Make a feature request!](https://feedback.mindee.com/?b=682f69c9e2404756e7e68d1c)

## Send a File and Poll

Use the client library to send the file to your OCR Model and return the result.

{% tabs %}
{% tab title="Python" %}
Requires Python ≥ 3.9. Python ≥ 3.11 is recommended.\
Requires the [Mindee Python SDK](https://pypi.org/project/mindee/) version **4.35.1** or greater.

{% code lineNumbers="true" %}

```python
from mindee import PathInput
from mindee.v2 import (
    Client,
    OCRParameters,
    OCRResponse,
)

input_path = "/path/to/the/file.ext"
api_key = "MY_API_KEY"
model_id = "MY_MODEL_ID"

# Init a new client
mindee_client = Client(api_key)

# Set Raw OCR parameters
model_params = OCRParameters(
    # ID of the model, required.
    model_id=model_id,
)

# Load a file from disk
input_source = PathInput(input_path)

# Send for processing using polling
response = mindee_client.enqueue_and_get_result(
    OCRResponse,
    input_source,
    model_params,
)

# Print a brief summary of the parsed data
print(response.inference)

# Access the OCR result
pages: list = response.inference.result.pages
```

{% endcode %}

Also take a look at the [OCR Result](https://docs.mindee.com/ocr-models/sdk-integration/ocr-result) documentation.
{% endtab %}

{% tab title="Node.js" %}
Requires Node.js ≥ 20.1. Node.js ≥ 22 is recommended.\
Requires the [Mindee Node.js SDK](https://www.npmjs.com/package/mindee/) version **5.5.0** or greater.

{% code lineNumbers="true" %}

```javascript
import * as mindee from "mindee";
// If you're on CommonJS:
// const mindee = require("mindee");

const apiKey = "MY_API_KEY";
const filePath = "/path/to/the/file.ext";
const modelId = "MY_MODEL_ID";

// Init a new client
const mindeeClient = new mindee.Client(
  { apiKey: apiKey }
);

// Set Raw OCR parameters
const modelParams = {
  modelId: modelId,
};

// Load a file from disk
const inputSource = new mindee.PathInput({ inputPath: filePath });

// Send for processing
const response = await mindeeClient.enqueueAndGetResult(
  mindee.product.Ocr,
  inputSource,
  modelParams,
);

// print a string summary
console.log(response.inference.toString());

// Access the result OCR pages
const pages = response.inference.result.pages;
```

{% endcode %}

Also take a look at the [OCR Result](https://docs.mindee.com/ocr-models/sdk-integration/ocr-result) documentation.
{% endtab %}

{% tab title="PHP" %}
Requires PHP ≥ 8.1. PHP ≥ 8.3 is recommended.\
Requires the [Mindee PHP SDK](https://packagist.org/packages/mindee/mindee) version **3.0.0** or greater.

{% code lineNumbers="true" %}

```php
<?php

use Mindee\Input\PathInput;
use Mindee\V2\Client;
use Mindee\V2\Product\Ocr\Params\OcrParameters;
use Mindee\V2\Product\Ocr\OcrResponse;

$apiKey = "MY_API_KEY";
$modelId = "MY_MODEL_ID";
$filePath = "/path/to/the/file.ext";

// Init a new client
$mindeeClient = new Client($apiKey);

// Set Raw OCR parameters
$modelParams = new OcrParameters(
    // ID of the model, required.
    $modelId,
);

// Load a file from disk
$inputSource = new PathInput($filePath);

// Send for processing using polling
$response = $mindeeClient->enqueueAndGetResult(
    OcrResponse::class,
    $inputSource,
    $modelParams
);

// Print a summary of the response
echo strval($response->inference);

// Access the ocr results
$pages = $response->inference->result->pages;
```

{% endcode %}

Also take a look at the [OCR Result](https://docs.mindee.com/ocr-models/sdk-integration/ocr-result) documentation.
{% endtab %}

{% tab title="Ruby" %}
Requires Ruby ≥ 3.2.\
Requires the [Mindee Ruby SDK](https://rubygems.org/gems/mindee) version **5.2.1** or greater.

{% code lineNumbers="true" %}

```ruby
require 'mindee'
require 'mindee/v2/product'

input_path = '/path/to/the/file.ext'
api_key = 'MY_API_KEY'
model_id = 'MY_MODEL_ID'

# Init a new client
mindee_client = Mindee::V2::Client.new(api_key: api_key)

# Set Raw OCR parameters
model_params = {
    # ID of the model, required.
    model_id: model_id,
}

# Load a file from disk
input_source = Mindee::Input::Source::PathInputSource.new(input_path)

# Send for processing
response = mindee_client.enqueue_and_get_result(
    Mindee::V2::Product::OCR::OCR,
    input_source,
    model_params
)

# Access the result OCR pages
puts response.inference.result.pages
```

{% endcode %}

Also take a look at the [OCR Result](https://docs.mindee.com/ocr-models/sdk-integration/ocr-result) documentation.
{% endtab %}

{% tab title="Java" %}
Requires Java ≥ 11. Java ≥ 17 is recommended.\
Requires the [Mindee Java SDK](https://central.sonatype.com/artifact/com.mindee.sdk/mindee-api-java) version **5.2.0** or greater.

{% code lineNumbers="true" %}

```java
import com.mindee.input.LocalInputSource;
import com.mindee.v2.MindeeClient;
import com.mindee.v2.product.ocr.OcrResponse;
import com.mindee.v2.product.ocr.params.OcrParameters;
import java.io.IOException;

public class SimpleMindeeClientV2 {

  public static void main(String[] args)
      throws IOException, InterruptedException
  {
    String apiKey = "MY_API_KEY";
    String filePath = "/path/to/the/file.ext";
    String modelId = "MY_MODEL_ID";

    // Init a new client
    var mindeeClient = new MindeeClient(apiKey);

    // Set Raw OCR parameters
    var modelParams = OcrParameters
        // ID of the model, required.
        .builder(modelId)
        .build();

    // Load a file from disk
    var inputSource = new LocalInputSource(filePath);

    // Send for processing using polling
    OcrResponse response = mindeeClient.enqueueAndGetResult(
        OcrResponse.class,
        inputSource,
        modelParams
    );

    // Print a summary of the response
    System.out.println(response.getInference().toString());

    // Access the result OCR pages
    var pages = response.getInference().getResult().getPages();
  }
}
```

{% endcode %}

Also take a look at the [OCR Result](https://docs.mindee.com/ocr-models/sdk-integration/ocr-result) documentation.
{% endtab %}

{% tab title=".NET" %}
.NET ≥ 8.0 is recommended.\
Requires the [Mindee .NET SDK](https://www.nuget.org/packages/Mindee) version **4.3.0** or greater.

{% code lineNumbers="true" %}

```csharp
using Mindee;
using Mindee.Input;
using Mindee.V2;
using Mindee.V2.Product.Ocr;
using Mindee.V2.Product.Ocr.Params;

string filePath = "/path/to/the/file.ext";
string apiKey = "MY_API_KEY";
string modelId = "MY_MODEL_ID";

// Construct a new client
Client mindeeClient = new Client(apiKey);

// Set Raw OCR parameters
var modelParams = new OcrParameters(
    modelId: modelId
);

// Load a file from disk
var inputSource = new LocalInputSource(filePath);

// Upload the file
var response = await mindeeClient.EnqueueAndGetResultAsync<OcrResponse>(
    inputSource, modelParams);

// Print a summary of the response
System.Console.WriteLine(response.Inference.ToString());

// Access the OCR result pages
var ocrPages = response.Inference.Result.Pages;
```

{% endcode %}

Also take a look at the [OCR Result](https://docs.mindee.com/ocr-models/sdk-integration/ocr-result) documentation.
{% endtab %}
{% endtabs %}


# OCR Result

Reference documentation on processing an OCR result using the Mindee SDKs.

You'll need to have a response object as described in the [Response Processing](/integrations/client-libraries-sdk/process-the-response) section.

## Accessing the Page OCR

An `OCRPage` describes the text and words of a single page in the document.

### `OCRPage` Attributes

#### Content

Full text content extracted from the document page.

#### Words

List of all words found on the page.

Each word has the following properties:

* Content: Text content of the word.
* Polygon: Coordinates of the detected word.

{% tabs %}
{% tab title="Python" %}

```python
from mindee import OCRResponse

def handle_response(response: OCRResponse) -> None:
    pages = response.inference.result.pages

    for page in pages:
        # Page-level properties
        page_content = page.content
        words = page.words

        # Access the full text content extracted from this page.
        print(f"Page content: {page_content}")

        # Access all words detected on this page.
        for word in words:
            word_content = word.content
            word_polygon = word.polygon

            print(f"Word: {word_content}")
            print(f"Polygon: {word_polygon}")
```

{% endtab %}

{% tab title="Node.js" %}

```javascript
handleResponse(response) {
  const pages = response.inference.result.pages;

  for (const page of pages) {
    // Full text content extracted from this page.
    const pageContent = page.content;
    console.log(`Page content: ${pageContent}`);

    // All words detected on this page.
    for (const word of page.words) {
      const wordContent = word.content;
      const wordPolygon = word.polygon;

      console.log(`Word: ${wordContent}`);
      console.log(`Polygon: ${wordPolygon.toString()}`);
    }
  }
}
```

{% endtab %}

{% tab title="PHP" %}

```php
use Mindee\V2\Product\Ocr\OcrResponse;

public function handleResponse(OcrResponse $response)
{
    $pages = $response->inference->result->pages;
}
```

{% endtab %}

{% tab title="Ruby" %}

```ruby
def handle_response(response)
  pages = response.inference.result.pages

  pages.each do |page|
    # Page-level properties
    page_content = page.content
    words = page.words

    # Access the full text content extracted from this page.
    puts "Page content: #{page_content}"

    # Access all words detected on this page.
    words.each do |word|
      word_content = word.content
      word_polygon = word.polygon

      puts "Word: #{word_content}"
      puts "Polygon: #{word_polygon}"
    end
  end
end
```

{% endtab %}

{% tab title="Java" %}

```java
import com.mindee.v2.product.ocr.OcrPage;
import com.mindee.v2.product.ocr.OcrResponse;
import com.mindee.v2.product.ocr.OcrWord;

public void handleResponse(OcrResponse response) {
    var pages = response.getInference().getResult().getPages();

    for (OcrPage page : pages) {
        // Page-level properties
        String pageContent = page.getContent();
        var words = page.getWords();

        // Access the full text content extracted from this page.
        System.out.println("Page content: " + pageContent);

        // Access all words detected on this page.
        for (OcrWord word : words) {
            String wordContent = word.getContent();
            var wordPolygon = word.getPolygon();

            System.out.println("Word: " + wordContent);
            System.out.println("Polygon: " + wordPolygon.toString());
        }
    }
}
```

{% endtab %}

{% tab title=".NET" %}

```csharp
using Mindee.V2.Product.Ocr;

public void HandleResponse(OcrResponse response)
{
    var pages = response.Inference.Result.Pages;

    foreach (var page in pages)
    {
        // Page-level properties
        string pageContent = page.Content;
        var words = page.Words;

        // Access the full text content extracted from this page.
        Console.WriteLine($"Page content: {pageContent}");

        // Access all words detected on this page.
        foreach (var word in words)
        {
            string wordContent = word.Content;
            var wordPolygon = word.Polygon;

            Console.WriteLine($"Word: {wordContent}");
            Console.WriteLine($"Polygon: {wordPolygon}");
        }
    }
}
```

{% endtab %}
{% endtabs %}


# No-Code Integration

No-Code and Low-Code integration support for Raw Text (OCR) models.

## Officially Supported Integrations

We're hard at work providing official integrations for Raw Text models on platforms that allow it.

Currently we have support for the following platforms:

* **n8n**: official support [coming soon](https://github.com/n8n-io/n8n/pull/18986)
* **Zapier**: officially supported
* **Make**: officially supported

Don't see support for your favorite platform? [Make a feature request!](https://feedback.mindee.com/?b=682f69c9e2404756e7e68d1c)

{% hint style="info" %}
**Usage is very similar to** [**Extraction models**](/extraction-models/no-code-integration)**.**

Complete documentation in progress...
{% endhint %}


# Integration Overview

Overview of connecting your system to the Mindee API.

## General Description

Mindee is ideal for handling large amounts of documents. The vast majority of our users will want to connect Mindee to their systems using our APIs.

It's possible to design any number of different use cases around document processing. These can be for purely internal processing or to provide final end users with a polished experience.

The API is asynchronous, RESTful, and returns objects.

## Before Starting

You'll need at least one model configured, see the [Models Overview](/models/models-overview) section for more details. This can be any model type (extraction or utility model), all models are integrated in a very similar way.

We recommend using the [Live Test](/models/live-test) feature before attempting to integrate the API.

You'll also need at least one API key, see the [Manage API Keys](/integrations/api-keys) section for more info.

## How to Integrate

All file processing routes are asynchronous, no synchronous routes are provided. You can either use a polling or webhook workflow.

If using our SDK integrations polling is abstracted away for you, meaning you can use a single synchronous method within the SDK to receive processing results.

These are are the fastest and easiest way to call our APIs, and allow our support teams to better help you.

### Client Libraries / SDKs

For a quick introduction and copy-paste ready code, look in the [Extraction Quick Start](/extraction-models/sdk-integration/quick-start) section.

{% hint style="success" %}
**Ask for Code Samples**

You can ask for specific code samples from the documentation AI.

Use the "Ask" button at the top of any page, or click below:

<button type="button" class="button primary" data-action="ask" data-query="List available model types and a brief description (Extraction, Split, Crop, Classify, OCR), allow to pick one. Then list available SDK languages and allow to pick one. Finally ask for model ID, and generate the code sample for polling." data-icon="gitbook-assistant">Ask "Write a code sample for me."</button>
{% endhint %}

Supported languages/frameworks: **Python**, **Node.js** (JS/TS), **PHP**, **Ruby**, **Java**, **.NET** (C#).

We provide full support for Client Libraries regardless of your plan. You can report any issues on our [bug tracker](https://feedback.mindee.com/?b=685c08afd7a1d2e47b124cbb) or directly on [GitHub](https://github.com/orgs/mindee/repositories).

### No-Code or Low-Code

If you're integrating using a no-code or low-code platform, take a look at the [No-Code Integration](/extraction-models/no-code-integration) section.

### Manual Integration

If none of the above options fit your requirements, take a look at the [Manual Integration](/integrations/api-reference) section.

{% hint style="warning" %}
**We do not recommend manually integrating**, and cannot guarantee full support.

Pro plans and above benefit from extended integration support.
{% endhint %}

## What to Send

You can send either a local file or an URL, it makes no difference for server-side processing.

However, when using our client libraries, you can [Load and Adjust a File](/integrations/client-libraries-sdk/load-and-adjust-a-file#adjust-the-source-file) if you have it locally.

## How to Receive Results

Inference operations are always asynchronous, meaning there is a route to POST the file and another mechanism to retrieve the results.

You can decide on using either the polling flow or the webhook flow.

[Using Polling](/integrations/polling-for-results) - uses a GET route, and is better suited for testing and small volumes.

[Using Webhooks](/integrations/webhooks) - sends directly to your server, and is more suited for heavy production use.

## Developing and Testing

A typical development release cycle could look like this:

1. Start with a new model. If you are making adjustments to an existing production model, we recommend copying it and only making changes to the copy.\
   Consider adding versioning info to the copied model's name, i.e. "Invoices v1.1" or "Receipts 2026-05-17".
2. Adjust your code as needed and test.
3. When deploying your code to testing or production environments, use the new model's ID.
4. After deployment, [lock your production model](/models/model-settings#locking-the-data-schema) to avoid accidental changes.
5. `GOTO 1`

## Frequently Asked Questions

<details>

<summary><strong>Can I use the same V1 API keys in the V2?</strong></summary>

No, V1 and V2 do not share API key information.

</details>

<details>

<summary><strong>Can I run V1 and V2 in parallel?</strong></summary>

Yes, absolutely.

Running both platforms in parallel is not only supported but recommended if you are migrating from V1 to V2.

All SDKs have full support for V1 and V2 running in parallel, more information here: [Client Libraries / SDKs](/integrations/client-libraries-sdk)

</details>

<details>

<summary><strong>Do you provide a testing or staging environment?</strong></summary>

We do not provide separate environments for development, staging, production, etc.

For information on testing your models and code before deploying to production, take a look at: [#developing-and-testing](#developing-and-testing "mention")

</details>

<details>

<summary><strong>How can I export the results to CSV?</strong></summary>

The Mindee API return contains lists of nested fields, it is an object-based format.\
The CSV format has no provisions for objects, it only has rows and columns.

To convert from Mindee's object format to a CSV format, you can either remove information or output to several CSV files (or sheets).

Since each model is different, and each use case is different, there is no universal way to do this conversion.

It's up to you to make a mapping that conforms to your needs. You can use the [SDKs](/extraction-models/sdk-integration) for easy manipulation of objects, or set up mapping rules in your [no-code](/extraction-models/no-code-integration) solution.

</details>


# Client Libraries / SDKs

Officially supported client libraries for integrating the Mindee platform.

By using the client libraries you'll be able to integrate faster and lower your maintenance costs.

Some useful tools are also provided, for example PDF processing and image compression.

You can use the client libraries to make API requests following the polling and webhook patterns.

All our client libraries are open-source (MIT license) and hosted on [GitHub](https://github.com/mindee).

Supported languages/frameworks: **Python**, **Node.js** (JS/TS), **PHP**, **Ruby**, **Java**, **.NET** (C#).

## Installation Instructions

{% tabs %}
{% tab title="Python" %}
Requires Python ≥ 3.9. Python ≥ 3.11 is recommended.

Simply install the [PyPi package](https://pypi.org/project/mindee/) using `pip`:

```sh
pip install -U mindee~=5.2
```

{% endtab %}

{% tab title="Node.js" %}
Requires Node.js ≥ 20.1. Node.js ≥ 22 is recommended.

Simply install the [NPM package](https://www.npmjs.com/package/mindee):

```sh
npm install mindee@^5.5.0
```

{% endtab %}

{% tab title="PHP" %}
Requires PHP ≥ 8.1. PHP ≥ 8.3 is recommended.

Simply install the [Packagist package](https://packagist.org/packages/mindee/mindee) using [composer](https://getcomposer.org/):

```sh
php composer.phar require "mindee/mindee:>=3.0"
```

{% endtab %}

{% tab title="Ruby" %}
Requires Ruby ≥ 3.2.

Simply install the [gem](https://rubygems.org/gems/mindee) using:

```shell
gem install mindee -v '~> 5.2'
```

{% endtab %}

{% tab title="Java" %}
Requires Java ≥ 11. Java ≥ 17 is recommended.

Group ID: `com.mindee.sdk`\
Artifact ID: `mindee-api-java`\
Version: `5.2.0` or greater

There are various installation methods, Maven, Gradle, etc:

[Installation Details](https://central.sonatype.com/artifact/com.mindee.sdk/mindee-api-java)
{% endtab %}

{% tab title=".NET" %}
.NET ≥ 8.0 is recommended.

Simply install the [NuGet package](https://www.nuget.org/packages/Mindee) using `dotnet add`:

```sh
dotnet add package Mindee --version 4.4
```

{% endtab %}
{% endtabs %}

Don't see support for your favorite language or framework? [Make a feature request!](https://feedback.mindee.com/?b=682f69c9e2404756e7e68d1c)

## Usage Details

Overall, the steps to using the Mindee service are:

1. [Client Configuration](/integrations/client-libraries-sdk/configure-the-client)
   1. Initialize the Mindee client.
   2. Set inference parameters, in particular the model ID to use.
2. [Load and Adjust a File](/integrations/client-libraries-sdk/load-and-adjust-a-file)
   1. Load a file from various supported sources: path, bytes, etc.
   2. *Optional*: adjust the source file before sending.
3. [Send a File or URL](/integrations/client-libraries-sdk/send-a-file-or-url)
   1. Send the file or an URL with the proper parameters.
4. [Response Processing](/integrations/client-libraries-sdk/process-the-response)
   1. Optional: load from a webhook.
   2. Optional: access document metadata
5. [Extraction Result](/extraction-models/sdk-integration/extraction-result)
   1. Handle the field values extracted from the document
   2. Optional: access field metadata (polygons, confidence score)

## Frequently Asked Questions

<details>

<summary><strong>Can I send requests in parallel?</strong></summary>

Yes. All clients can be used to send requests in parallel.

The exact implementation is left to the user:

* For Node.js, you'll want to use asynchronous processing.
* For all others, you'll want to use threads or processes.

</details>

<details>

<summary><strong>Can I use the v1 and v2 APIs together?</strong></summary>

Yes. Each client library has support for both v1 and v2 APIs.

You'll need to make a separate instance of the client classes:

* .NET and Java, use `MindeeClient` and `MindeeClientV2`
* all others, use `Client` and `ClientV2`

The code to make requests and to process results is **very different** between v1 and v2.

We highly recommend having different files (or even modules) for handling each API version.

</details>

<details>

<summary><strong>Can I send only a specific page of a multi-page PDF?</strong></summary>

Yes. All libraries have support for cutting/extracting PDF pages.

For more information, consult: [Load and Adjust a File](/integrations/client-libraries-sdk/load-and-adjust-a-file#manipulate-pdf-pages).

</details>

<details>

<summary><strong>How to stop PDFs with too many pages from being sent?</strong></summary>

**Do not use file size**, a text PDF with 200 pages can be smaller than a single photo.

Much more reliable to count the actual number of pages in the PDF document.

Use the built-in file metadata methods and properties to easily add business rules based on the number of pages (among other data).

For more information, consult: [Load and Adjust a File](/integrations/client-libraries-sdk/load-and-adjust-a-file#source-file-metadata).

</details>

<details>

<summary><strong>I'm using a Supabase edge function, should I use the API directly?</strong></summary>

We recommend using the Mindee [Node.js client library](https://github.com/mindee/mindee-api-nodejs) in Supabase.

You can install it in your edge function(s) using `npm`.

</details>

<details>

<summary><strong>Which library features are officially supported?</strong></summary>

Anything documented here is officially supported and is considered stable for production use.

Anything in a library that is not documented here, is **not** officially supported and subject to change or removal.

</details>

<details>

<summary><strong>There is a bug with the SDK, how can I get help?</strong></summary>

If you are encountering a persistant issue, you should try first:

1. Asking the Docs Assistant for help: <button type="button" class="button primary" data-action="ask" data-icon="gitbook-assistant">Ask a question...</button>
2. Asking your local agent for help by using one of our [SKILL files](/integrations/ai-coding-assistants#skill-files)

After this, if you determine (or are told) there is a problem with the SDK itself, the best bet would be to open a bug report on the relevant SDK's GitHub repository:

* **Python**: <https://github.com/mindee/mindee-api-python/issues>
* **Node.js**: <https://github.com/mindee/mindee-api-nodejs/issues>
* **PHP**: <https://github.com/mindee/mindee-api-php/issues>
* **Ruby**: <https://github.com/mindee/mindee-api-ruby/issues>
* **Java**: <https://github.com/mindee/mindee-api-java/issues>
* **.NET**: <https://github.com/mindee/mindee-api-dotnet/issues>

Our SDK engineers will be automatically notified of the new issue.

**Note:** Pro and above plans can contact our support teams directly.

</details>


# Command Line Tools (CLI)

For quick testing you can use the integrated CLI tools that ship with the Client Libraries.

All our APIs are asynchronous, as a result it's not very practical to use cURL or equivalent for quick testing.

For this reason the SDKs include a Command Line Interface (CLI) to allow for quick testing and debugging of your models.

{% tabs %}
{% tab title="Python" %}

```bash
# Install the library
pip install mindee

# General help
mindee --help

# Help for extraction models
mindee extraction --help

# Run extraction on a file
#   -k: API key, you can instead `export MINDEE_V2_API_KEY=md_XXXXXXXXXXXX`
#   -m: Your model ID
#   Positional arg: Absolute path to the file
mindee extraction \
  -k md_XXXXXXXXXXXX \
  -m xxxx-xxxx-xxxx-xxxx-xxxx \
  /path/to/the/file.pdf
```

{% endtab %}

{% tab title="Node.js" %}

```shellscript
# Install the library
npm install mindee

# General help
./node_modules/.bin/mindeeV2 --help

# Help for extraction models
./node_modules/.bin/mindeeV2 extraction --help

# Run extraction on a file
#   -k: API key, you can instead `export MINDEE_V2_API_KEY=md_XXXXXXXXXXXX`
#   -m: Your model ID
#   Positional arg: Absolute path to the file
./node_modules/.bin/mindeeV2 extraction \
  -k md_XXXXXXXXXXXX \
  -m xxxx-xxxx-xxxx-xxxx-xxxx \
  /path/to/the/file.pdf
```

{% endtab %}

{% tab title="Ruby" %}

```bash
# General help
./bin/mindee.rb --help

# Help for extraction models
./bin/mindee.rb extraction --help

# Run extraction on a file
#   -k: API key, you can instead `export MINDEE_V2_API_KEY=md_XXXXXXXXXXXX`
#   -m: Your model ID
#   Positional arg: Absolute path to the file
./bin/mindee.rb extraction \
  -k md_XXXXXXXXXXXX \
  -m xxxx-xxxx-xxxx-xxxx-xxxx \
  /path/to/the/file.pdf
```

{% endtab %}

{% tab title="PHP" %}

```bash
# General help
php ./bin/cli.php --help

# Help for extraction models
php ./bin/cli.php extraction --help

# Run extraction on a file
#   -k: API key, you can instead `export MINDEE_V2_API_KEY=md_XXXXXXXXXXXX`
#   -m: Your model ID
#   Positional arg: Absolute path to the file
php ./bin/cli.php extraction \
  -k md_XXXXXXXXXXXX \
  -m xxxx-xxxx-xxxx-xxxx-xxxx \
  /path/to/the/file.pdf
```

{% endtab %}

{% tab title="Java" %}
Use from the root of the [Git repository](https://github.com/mindee/mindee-api-java).

```bash
# Compile locally
mvn clean verify

# General help
./cli.sh --help

# Help for extraction models
./cli.sh extraction --help

# Run extraction on a file
#   -k: API key, you can instead `export MINDEE_V2_API_KEY=md_XXXXXXXXXXXX`
#   -m: Your model ID
#   Positional arg: Absolute path to the file
./cli.sh extraction \
  -k md_XXXXXXXXXXXX \
  -m xxxx-xxxx-xxxx-xxxx-xxxx \
  /path/to/the/file.pdf
```

{% endtab %}

{% tab title=".NET" %}
First install the CLI package: <https://www.nuget.org/packages/Mindee.Cli>

Note: this is a separate package from the general Mindee SDK.

```bash
# Install the package
dotnet tool install Mindee.Cli

# General help
dotnet mindee --help

# Help for extraction models
dotnet mindee extraction --help

# Run extraction on a file
#   -k: API key, you can instead `export MINDEE_V2_API_KEY=md_XXXXXXXXXXXX`
#   -m: Your model ID
#   Positional arg: Absolute path to the file
dotnet mindee extraction \
  -k md_XXXXXXXXXXXX \
  -m xxxx-xxxx-xxxx-xxxx-xxxx \
  /path/to/the/file.pdf
```

{% endtab %}
{% endtabs %}

{% hint style="info" %}
We are still working hard to make the CLI utilities more complete and useful.\
**As a result, the interface is not yet considered stable, and may be subject to change.**

If you have any ideas to improve them, [make a feature request!](https://feedback.mindee.com/?b=682f69c9e2404756e7e68d1c)
{% endhint %}


# Client Configuration

Reference documentation on preparing and configuring the Mindee client.

{% hint style="info" %}
**This is reference documentation.**

Code samples shown are only examples, and will not work as-is.\
You'll need to copy-paste and modify according to your requirements.

Looking full code samples?

\ <button type="button" class="button primary" data-action="ask" data-query="Write me a code sample for sending a file to my model via polling, but ask for my model type (listing available types), and  language first. Assume &#x22;MY_MODEL_ID&#x22; for the model ID parameter. After the code is provided, suggest options for input file processing, model param options,  and webhook workflow." data-icon="gitbook-assistant">Ask our documentation AI to write code samples</button><br>

You can also use the "Ask" button at the top of any page in the documentation.
{% endhint %}

## Before You Start

You'll need to have one of the [official Mindee SDKs](/integrations/client-libraries-sdk) installed.

You'll also need to use one of your [API keys](/integrations/api-keys) and have at least one [Model](/models/models-overview) configured.

The client works for all model types.

## Initialize the Mindee Client

Before sending any files to the Mindee servers for processing, you'll need to initialize your client.

This should be the first step in your code. It will determine which organization is used to make the calls.

You should reuse the same client instance for all calls of the same organization.

The client instance is thread-safe in applicable languages.

{% tabs %}
{% tab title="Python" %}
First import the needed classes:

```python
from mindee.v2 import Client
```

For the API key, you can pass it directly to the client.\
This is useful for quick testing.

```python
api_key = "MY_API_KEY"

mindee_client = Client(api_key)
```

Instead of passing the key directly, you can also set the following environment variable:

`MINDEE_V2_API_KEY`

This is recommended for production use.\
In this way there is no need to pass the `api_key` when initializing the client.

```python
mindee_client = Client()
```

{% endtab %}

{% tab title="Node.js" %}
First import the needed classes. We recommend using ES Modules.

```typescript
import * as mindee from "mindee";
// If you're on CommonJS:
// const mindee = require("mindee");
```

For the API key, you can pass it directly to the client.\
This is useful for quick testing.

```typescript
const apiKey = "MY_API_KEY";

const mindeeClient = new mindee.Client({ apiKey: apiKey });
```

Instead of passing the key directly, you can also set the following environment variable:

`MINDEE_V2_API_KEY`

This is recommended for production use.\
In this way there is no need to pass the `apiKey` argument when initializing the client.

```typescript
const mindeeClient = new mindee.Client();
```

**Advanced Usage**

internally, [undici](https://undici.nodejs.org/) is used for making HTTP calls, and when the client is initialized `getGlobalDispatcher` is called.

**This is perfectly fine for the vast majority of cases:**\
If you set a custom Agent as your global dispatcher, Mindee will use it for all calls.

In some rare cases you may need a specific dispatcher for Mindee, different from your global dispatcher. You can set a custom dispatcher as follows:

```javascript
import { Agent, interceptors } from "undici";

// Init the custom Dispatcher
// REPLACE THE LINE BELOW WITH YOUR CODE - this example will NOT work!
const myDispatcher = new Agent().compose( ... );

const mindeeClient = new mindee.Client({
  // Don't set if using an environment variable
  apiKey: apiKey,

  // Activates verbose logging - disable in production!
  debug: true,

  // Pass the custom dispatcher to the Mindee client
  dispatcher: myDispatcher,
});
```

The Mindee client will then use the given dispatcher for **all calls**.

This is the recommended setup when needing to pass extra parameters or configuration options, for example when using the Client behind a proxy.

When running in some serverless environments like Vercel, *sometimes* (this depends on your exact setup) there are errors caused by freezing/thawing of the process. These can be avoided by:

<pre class="language-javascript"><code class="lang-javascript"><strong>// Avoids stale connections accumulating in the pool.
</strong>const myDispatcher = new Agent({
  // Drop idle sockets after 1 ms, effectively disables connection reuse
  keepAliveTimeout: 1,
  // Prevent the server from extending the window via a Keep-Alive header
  keepAliveMaxTimeout: 1,
});
</code></pre>

{% endtab %}

{% tab title="PHP" %}
First import the needed classes:

```php
use Mindee\V2\Client;
```

For the API key, you can pass it directly to the client.\
This is useful for quick testing.

```php
$apiKey = "MY_API_KEY";

$mindeeClient = new Client($apiKey);
```

Instead of passing the key directly, you can also set the following environment variable:

`MINDEE_V2_API_KEY`

This is recommended for production use.\
In this way there is no need to pass the `$apiKey` when initializing the client.

```php
$mindeeClient = new Client($apiKey);
```

{% endtab %}

{% tab title="Ruby" %}
First import the Mindee package:

```ruby
require 'mindee'
```

For the API key, you can pass it directly to the client.\
This is useful for quick testing.

```ruby
api_key = 'MY_API_KEY'
mindee_client = Mindee::V2::Client.new(api_key: api_key)
```

Instead of passing the key directly, you can also set the following environment variable:

`MINDEE_V2_API_KEY`

This is recommended for production use.\
In this way there is no need to pass the `api_key` when initializing the client.

```ruby
mindee_client = Mindee::V2::Client.new()
```

{% endtab %}

{% tab title="Java" %}
First import the needed classes:

```java
import com.mindee.v2.MindeeClient;
```

For the API key, you can pass it directly to the client.\
This is useful for quick testing.

```java
String apiKey = "MY_API_KEY";

MindeeClient mindeeClient = new MindeeClient(apiKey);
```

Instead of passing the key directly, you can also set the following environment variable:

`MINDEE_V2_API_KEY`

This is recommended for production use.\
In this way there is no need to pass the `apiKey` argument when initializing the client.

```java
MindeeClient mindeeClient = new MindeeClient();
```

{% endtab %}

{% tab title=".NET" %}
First add the required namespaces.

```csharp
using Mindee.V2;
```

For the API key, you can pass it directly to the client.\
This is useful for quick testing.

```csharp
string apiKey = "MY_API_KEY";

Client mindeeClient = new Client(apiKey);
```

Instead of passing the key directly, you can also set the following environment variable:

`MindeeV2__ApiKey`

This is recommended for production use.\
In this way there is no need to pass the `apiKey` argument when initializing the client.

```csharp
Client mindeeClient = new Client();
```

{% endtab %}
{% endtabs %}


# Basic Model Configuration

Reference documentation on preparing and configuring the Mindee model inference.

{% hint style="info" %}
**This is reference documentation.**

Code samples shown are only examples, and will not work as-is.\
You'll need to copy-paste and modify according to your requirements.

Looking full code samples?

\ <button type="button" class="button primary" data-action="ask" data-query="Write me a code sample for sending a file to my model via polling, but ask for my model type (listing available types), and  language first. Assume &#x22;MY_MODEL_ID&#x22; for the model ID parameter. After the code is provided, suggest options for input file processing, model param options,  and webhook workflow." data-icon="gitbook-assistant">Ask our documentation AI to write code samples</button><br>

You can also use the "Ask" button at the top of any page in the documentation.
{% endhint %}

Parameters that apply to all Mindee model types (Extraction, Split, Crop, etc).

Inference parameters control:

* which model to use
* server-side processing options

All model parameter classes inherit from a base class. In these samples we will be showing the general usage and parameters common to all.

All examples use the Extraction model (product in the SDK), but you'll need to adjust depending on the model type you are using.

The available classes are:

* `ExtractionParameters`  for [Extraction](/extraction-models/extraction-models-overview) models
* `SplitParameters`  for [Split](/split-models/split) models
* `CropParameters`  for [Crop](/crop-models/crop) models
* `ClassificationParameters`  for [Classification](/classification-models/classification) models
* `OCRParameters`  or `OcrParameters`  for [Raw Text OCR](/raw-text-ocr-models/ocr) models

### Use an Alias

The optional `alias` argument lets you attach your own identifier to a request as a free-form string.

This alias could be an internal document ID, a reference number, or a database key.

It is echoed back unchanged in both the job and result responses, making it straightforward to match API responses with your own records.

Aliases are not unique in Mindee, you can use the same alias value multiple times.

{% tabs %}
{% tab title="Python" %}

```python
inference_params = ExtractionParameters(
    # ID of the model, required.
    model_id="MY_MODEL_ID",
    
    # Use an alias to link the file to your own DB.
    # If set, it will be included in the job and result responses.
    alias="internal-doc-id-123",
    
    # ... any other options ...
)
```

{% endtab %}

{% tab title="Node.js" %}

```typescript
const modelParams = {
  // ID of the model, required.
  modelId: "MY_MODEL_ID",
  
  // Use an alias to link the file to your own DB.
  // If set, it will be included in the job and result responses.
  alias: "internal-doc-id-123",
  
  // ... any other options ...
};
```

{% endtab %}

{% tab title="PHP" %}

```php
$modelParams = new ExtractionParameters(
    // ID of the model, required.
    "MY_MODEL_ID",

    // Use an alias to link the file to your own DB.
    // If set, it will be included in the job and result responses.
    alias: "internal-doc-id-123",
    
    // ... any other options ...
);
```

{% endtab %}

{% tab title="Ruby" %}

```ruby
inference_params = {
    # ID of the model, required.
    model_id: 'MY_MODEL_ID',
    
    # Use an alias to link the file to your own DB.
    # If set, it will be included in the job and result responses.
    file_alias: 'internal-doc-id-123',
    
    # ... any other options ...
}
```

{% endtab %}

{% tab title="Java" %}

```java
var modelParams = ExtractionParameters
    // ID of the model, required.
    .builder("MY_MODEL_ID")
    
    // Use an alias to link the file to your own DB.
    // If set, it will be included in the job and result responses.
    .alias("internal-doc-id-123")
    
    // ... any other options ...

    // complete the builder
    .build();
```

{% endtab %}

{% tab title=".NET" %}

```csharp
var modelParams = new ExtractionParameters(
    // ID of the model, required.
    modelId: "MY_MODEL_ID"

    // Use an alias to link the file to your own DB.
    // If set, it will be included in the job and result responses.
    , alias: "internal-doc-id-123"
    
    // ... any other options ...
);
```

{% endtab %}
{% endtabs %}


# Load and Adjust a File

Reference documentation on loading and manipulating files prior to sending, using Mindee client libraries.

{% hint style="info" %}
**This is reference documentation.**

Code samples shown are only examples, and will not work as-is.\
You'll need to copy-paste and modify according to your requirements.

Looking full code samples?

\ <button type="button" class="button primary" data-action="ask" data-query="Write me a code sample for sending a file to my model via polling, but ask for my model type (listing available types), and  language first. Assume &#x22;MY_MODEL_ID&#x22; for the model ID parameter. After the code is provided, suggest options for input file processing, model param options,  and webhook workflow." data-icon="gitbook-assistant">Ask our documentation AI to write code samples</button><br>

You can also use the "Ask" button at the top of any page in the documentation.
{% endhint %}

## Before Starting

In most cases you'll be loading a source file for use in the Mindee Client.

Initialize your client as described in the section: [Client Configuration](/integrations/client-libraries-sdk/configure-the-client).

However, you don't actually need the client initialized to use these features, only the client library installed.

## Overview

Overall, the steps to sending a file are:

1. Load a source file.
2. *Optional*: adjust the source file before sending.
3. Use the Mindee client instance to send the file.

## Load a Source File

You can load a source file from a path, from raw bytes, from a bytes stream, or from a language-specific object. Choose the appropriate type based on your application requirements.

If you're unsure of which to use, we recommend loading from a path.

{% tabs %}
{% tab title="Python" %}
To load a file, you'll need to import the corresponding input class from the `mindee` module.

To load a path string, use `PathInput` .

```python
from mindee import PathInput

input_path = "/path/to/the/file.ext"
input_source = PathInput(input_path)
```

To load a `Path` instance, use `PathInput`.

```python
from pathlib import Path
from mindee import PathInput

input_path = Path("/path/to/the/file.ext")
input_source = PathInput(input_path)
```

To load raw bytes, use `BytesInput` . The `filename` parameter is required.

```python
from pathlib import Path
from mindee import BytesInput

input_path = Path("/path/to/the/file.ext")
with input_path.open("rb") as fh:
    input_bytes = fh.read()

input_source = BytesInput(
    input_bytes,
    filename="file.ext",
)
```

To load a base-64 string, use `Base64Input` . The `filename` parameter is required.\
The string will be decoded into bytes internally.

```python
from mindee import Base64Input

input_base64 = "iVBORw0KGgoAAAANSUhEUgAAABgAAA ..."
input_source = Base64Input(
    input_base64,
    filename="file.ext",
)
```

To load a file handle, use `FileInput`.\
It **must** be opened in binary mode, as a `BinaryIO` .

```python
from pathlib import Path
from mindee import FileInput

input_path = Path("/path/to/the/file.ext")
with input_path.open("rb") as fh:
    input_source = FileInput(fh)
    # IMPORTANT:
    # Continue all operations inside the 'with' statement.
    response = mindee_client.enqueue_and_get_result(
        InferenceResponse,
        input_source,
        params,
    )
```

{% endtab %}

{% tab title="Node.js" %}
To load a file, you'll need to import the corresponding input class and instantiate it.

Make sure to import the needed classes:

```typescript
import * as mindee from "mindee";
// If you're on CommonJS:
// const mindee = require("mindee");
```

To load a path string, use `PathInput`.

```typescript
const filePath = "/path/to/the/file.ext";
const inputSource = new mindee.PathInput({ inputPath: filePath });
```

To load a `Buffer` instance, use `BufferInput` . The `filename` parameter is required.

```typescript
const buffer = Buffer.from(
  await fs.promises.readFile("/path/to/the/file.ext")
);

const inputSource = new mindee.BufferInput({
  buffer: buffer,
  filename: "file.ext",
});
```

To load raw bytes, use `BytesInput` . The `filename` parameter is required.

```typescript
const inputBytes = new Uint8Array(
  await fs.promises.readFile("/path/to/the/file.ext")
);

const inputSource = new mindee.BytesInput({
  inputBytes: inputBytes,
  filename: "file.ext",
});
```

To load a `Stream`, use `StreamInput`. The `filename` parameter is required.

```typescript
const stream = fs.createReadStream("/path/to/the/file.ext");

const inputSource = new mindee.StreamInput({
  inputStream: stream,
  filename: "file.ext",
});
```

To load a base-64 string, use `Base64Input` . The `filename` parameter is required.

```typescript
const b64String = "iVBORw0KGgoAAAANSUhEUgAAABgAAA ...";

const inputSource = new mindee.Base64Input({
  inputString: b64String,
  filename: "file.ext",
});
```

{% endtab %}

{% tab title="PHP" %}
To load a file, you'll need to import the corresponding input class from the `Mindee\Input` namespace.

To load a path string, use `PathInput`.

```php
use Mindee\Input\PathInput;

$filePath = "/path/to/the/file.ext";
$inputSource = new PathInput($filePath);
```

To load a file resource, use `FileInput`.

```php
use Mindee\Input\FileInput;

$handle = fopen("/path/to/the/file.ext", "rb");
$inputSource = new FileInput($handle);
```

To load raw bytes, use `BytesInput`. The filename is required.

```php
use Mindee\Input\BytesInput;

$filePath = "/path/to/the/file.ext";
$handle = fopen($filePath, "rb");
$contents = fread($handle, filesize($filePath));

$inputSource = new BytesInput($contents, "file.ext");
```

To load a base-64 string, use `Base64Input` .\
The string will be decoded into bytes internally. The filename is required.

```php
use Mindee\Input\Base64Input;

$inputBase64 = "iVBORw0KGgoAAAANSUhEUgAAABgAAA ..."

$inputSource = Base64Input($inputBase64, "file.ext");
```

{% endtab %}

{% tab title="Ruby" %}
To load a path string, use the `PathInputSource` class.

```ruby
input_path = '/path/to/the/file.ext'
input_source = Mindee::Input::Source::PathInputSource.new(input_path)
```

To load raw bytes, use the `BytesInputSource` class.

```ruby
input_bytes = File.binread('/path/to/the/file.ext')
input_source = Mindee::Input::Source::BytesInputSource.new(input_bytes, file_name)
```

To load a base-64 string, use `Base64InputSource`. The filename is required.\
The string will be decoded into bytes internally.

```ruby
input_base64 = 'iVBORw0KGgoAAAANSUhEUgAAABgAAA ...'
input_source = Mindee::Input::Source::Base64InputSource.new(
    input_base64, 'file.ext'
)
```

To load a file handle, use `FileInputSource`. The filename is required.\
It must be opened in binary mode.

```ruby
file = File.open('/path/to/the/file.ext', 'rb')
input_source = Mindee::Input::Source::FileInputSource.new(file, 'file.ext')
```

{% endtab %}

{% tab title="Java" %}
To load a file, initialize it using the `LocalInputSource` class.

This class has different constructors to allow for opening various types of inputs.

To load a path string:

```java
var filePath = "/path/to/the/file.ext";

var inputSource = new LocalInputSource(filePath);
```

To load a `Path` instance:

```java
var filePath = new Path("/path/to/the/file.ext");

var inputSource = new LocalInputSource(filePath);
```

To load a `File` instance:

```java
var file = new File("/path/to/the/file.ext");

var inputSource = new LocalInputSource(file);
```

To load a byte array, the filename is required:

```java
var fileBytes = Files.readAllBytes("/path/to/the/file.ext");
var filename = "file.ext";

var inputSource = new LocalInputSource(fileBytes, filename);
```

To load an `InputStream` instance, the filename is required:

```java
var fileStream = new FileInputStream(
    new File("/path/to/the/file.ext")
);
var filename = "file.ext";

var inputSource = new LocalInputSource(fileStream, filename);
```

To load a base-64 string, the filename is required:

```java
var inputBase64 = "iVBORw0KGgoAAAANSUhEUgAAABgAAA ...";
var filename = "file.ext";

var inputSource = new LocalInputSource(inputBase64, filename);
```

{% endtab %}

{% tab title=".NET" %}
To load a file, initialize it using the `LocalInputSource` class.

This class has different constructors to allow for opening various types of inputs.

To load a path string:

```csharp
string filePath = "/path/to/the/file.ext";

var inputSource = new LocalInputSource(filePath)
```

To load a `FileInfo` instance:

```csharp
FileInfo fileinfo = new FileInfo("/path/to/the/file.ext");

var inputSource = new LocalInputSource(fileinfo)
```

To load a byte array, the filename is required:

```csharp
byte[] fileBytes = File.ReadAllBytes("/path/to/the/file.ext");
string filename = "file.ext";

var inputSource = new LocalInputSource(fileBytes, filename)
```

To load a `Stream` instance, the filename is required:

```csharp
Stream fileStream = FileStream SourceStream = File.Open(
    "/path/to/the/file.ext", FileMode.Open);
string filename = "file.ext";

var inputSource = new LocalInputSource(fileStream, filename)
```

{% endtab %}
{% endtabs %}

## Source File Metadata

Once a source file is loaded, various metadata can be accessed.

These metadata are overall very similar to the ones returned by the API after a successful inference.\
If you have need of these metadata before getting the full result back from the server, this is the recommended way.

This can be useful for applying business rules based on the input file, for example:

* Send PDFs to one model, images to another
* Don't send PDFs with too many pages
* Save the filename to a database
* ...

Here are some code samples, using an [*input source*](/integrations/client-libraries-sdk/load-and-adjust-a-file#load-a-source-file) instance.

{% tabs %}
{% tab title="Python" %}

```python
filename: str = input_source.filename
is_pdf: bool = input_source.is_pdf
number_of_pages: int = input_source.page_count
```

{% endtab %}

{% tab title="Node.js" %}

```javascript
// make sure to initialze the source first
await inputSource.init();

const filename = inputSource.filename;
const isPdf = inputSource.isPdf();
const numberOfPages = await inputSource.getPageCount();
```

{% endtab %}

{% tab title="PHP" %}

```php
$filename = $inputSource->fileName;
$isPdf = $inputSource->isPdf();
$numberOfPages = $inputSource->getPageCount();
```

{% endtab %}

{% tab title="Ruby" %}

```ruby
filename = input_source.filename
is_pdf = input_source.pdf?
number_of_pages = input_source.page_count
```

{% endtab %}

{% tab title="Java" %}

```java
String filename = inputSource.getFilename();
boolean isPdf = inputSource.isPdf();
int numberOfPages = inputSource.getPageCount();
```

{% endtab %}

{% tab title=".NET" %}

```csharp
string filename = inputSource.Filename;
bool isPdf = inputSource.IsPdf();
int numberOfPages = inputSource.GetPageCount();
```

{% endtab %}
{% endtabs %}

## Adjust the Source File

Optionally make changes and adjustments to the source file before sending.

{% hint style="info" %}
All file adjustments are applied in-memory to the source file instance.

If loaded from disk, the original file is not modified.
{% endhint %}

### Fix PDF Headers

In some cases, PDFs will have corrupt or invalid headers.\
These files will return a 4xx HTTP error as the server will be unable to process them.

You can try to fix the headers using the provided functions.

**Note:** this feature is not yet available for all languages.

Here are some code samples, using an [*input source*](/integrations/client-libraries-sdk/load-and-adjust-a-file#load-a-source-file) instance.

{% tabs %}
{% tab title="Python" %}

```python
input_source.fix_pdf()
```

{% endtab %}

{% tab title="PHP" %}

```php
$inputSource->fixPDF();
```

{% endtab %}

{% tab title="Ruby" %}

```ruby
input_source.fix_pdf!
```

{% endtab %}
{% endtabs %}

### Compress Files

There is no need to send excessively large files to the Mindee API.

Unfortunately, many modern smartphones can take very high resolution images.

We provide a way to compress images before sending to the API.

Here are some code samples, using an [*input source*](/integrations/client-libraries-sdk/load-and-adjust-a-file#load-a-source-file) instance.

{% tabs %}
{% tab title="Python" %}
Basic usage is very simple, and can be applied to both images and PDFs:

```python
input_source.compress(quality=85)
```

For images, you can also set a maximum height and/or width.\
The aspect ratio will always be preserved.

For example to compress and resize to no greater than 1920x1920 pixels:

```python
input_source.compress(
    quality=85, max_width=1920, max_height=1920
)
```

{% endtab %}

{% tab title="Node.js" %}
Basic usage is very simple, and can be applied to both images and PDFs:

```typescript
await inputSource.compress(85);
```

For images, you can also set a maximum height and/or width.\
The aspect ratio will always be preserved.

For example to compress and resize to no greater than 1920x1920 pixels:

```typescript
await inputSource.compress(85, 1920, 1920);
```

{% endtab %}

{% tab title="PHP" %}
Basic usage is very simple, and can be applied to both images and PDFs:

```php
$inputSource->compress(quality: 85);
```

For images, you can also set a maximum height and/or width.\
The aspect ratio will always be preserved.

For example to compress and resize to no greater than 1920x1920 pixels:

```php
$inputSource->compress(
    quality: 85, maxWidth: 1920, maxHeight: 1920
);
```

{% endtab %}

{% tab title="Ruby" %}
Basic usage is very simple, and can be applied to both images and PDFs:

```ruby
input_source.compress!(quality:85)
```

For images, you can also set a maximum height and/or width. The aspect ratio will always be preserved.\
For example to compress and resize to no greater than 1920x1920 pixels:

```ruby
input_source.compress!(
    quality:85, max_width:1920, max_height:1920
)
```

{% endtab %}

{% tab title="Java" %}
Basic usage is very simple, and can be applied to both images and PDFs:

```java
inputSource.compress(85);
```

For images, you can also set a maximum height and/or width.\
The aspect ratio will always be preserved.

For example to compress and resize to no greater than 1920x1920 pixels:

```java
inputSource.compress(85, 1920, 1920);
```

{% endtab %}

{% tab title=".NET" %}
Basic usage is very simple, and can be applied to both images and PDFs:

```csharp
inputSource.Compress(quality: 85);
```

For images, you can also set a maximum height and/or width.\
The aspect ratio will always be preserved.

For example to compress and resize to no greater than 1920x1920 pixels:

```csharp
inputSource.Compress(
    quality: 85, maxWidth: 1920, maxHeight: 1920);
```

{% endtab %}
{% endtabs %}

### Manipulate PDF Pages

In some cases, PDFs will have some superfluous pages present.

For example a cover page or terms and conditions which are not useful to the desired data extraction.

These extra pages count towards your billing and slow down processing.

It is therefore in your best interest to remove them before sending.

**Parameters:**

* "Page Indexes" is required and is a list of 0-based page indexes.\
  Use negative values to specify indexes starting from the end, i.e. `-1` for the last page.
* "Operation" specifies whether to keep only specified pages or remove specified pages.\
  One of "Keep Only" or "Remove".
* "On Min Pages" is optional and specifies the minimum number of pages a document must have for the operation to take place. The value of `0` means any number of pages.

Exact naming of parameters will depend on the language.

Here are some code samples, using an [*input source*](/integrations/client-libraries-sdk/load-and-adjust-a-file#load-a-source-file) instance.

{% tabs %}
{% tab title="Python" %}

```python
from mindee import PageOptions

# Set the options as follows:
# For all documents, keep only the first page
page_options = PageOptions(
    operation="KEEP_ONLY",
    page_indexes=[0],
)

# Apply in-memory
input_source.apply_page_options(page_options)
```

Some other examples:

```python
# Only for documents having 3 or more pages:
# Keep only these pages: first, penultimate, last
PageOptions(
    operation="KEEP_ONLY",
    on_min_pages=3,
    page_indexes=[0, -2, -1],
)

# For all documents:
# Remove the first page
PageOptions(
    operation="REMOVE",
    page_indexes=[0],
)

# Only for documents having 10 or more pages:
# Remove the first 5 pages
PageOptions(
    operation="REMOVE",
    on_min_pages=10,
    page_indexes=list(range(5)),
)
```

{% endtab %}

{% tab title="Node.js" %}

```typescript
// Set the options as follows:
// For all documents, keep only the first page
const pageOptions: mindee.PageOptions = {
  operation: mindee.PageOptionsOperation.KeepOnly,
  pageIndexes: [0],
};

// Apply in-memory
await inputSource.applyPageOptions(pageOptions);
```

Some other examples:

```typescript
// Only for documents having 3 or more pages:
// Keep only these pages: first, penultimate, last
const pageOptions: mindee.PageOptions = {
  operation: mindee.PageOptionsOperation.KeepOnly,
  onMinPages: 3,
  pageIndexes: [0, -2, -1],
};

// For all documents:
// Remove the first page
const pageOptions: mindee.PageOptions = {
  operation: mindee.PageOptionsOperation.Remove,
  pageIndexes: [0],
};

// Only for documents having 10 or more pages:
// Remove the first 5 pages
const pageOptions: mindee.PageOptions = {
  operation: mindee.PageOptionsOperation.Remove,
  onMinPages: 10,
  pageIndexes: [0, 1, 2, 3, 4],
};
```

{% endtab %}

{% tab title="PHP" %}

```php
use Mindee\Input\PageOptions;
use const Mindee\Input\KEEP_ONLY;

// Set the options as follows:
// For all documents, keep only the first page
$pageOptions = new PageOptions(
    pageIndexes: [0],
    operation: KEEP_ONLY
);

// Apply in-memory
$inputSource->applyPageOptions($pageOptions);
```

Some other examples:

```php
use Mindee\Input\PageOptions;
use const Mindee\Input\KEEP_ONLY;
use const Mindee\Input\REMOVE;

// Only for documents having 3 or more pages:
// Keep only these pages: first, penultimate, last
$pageOptions = new PageOptions(
    pageIndexes: [0, -2, -1],
    operation: KEEP_ONLY,
    onMinPage: 3
);

// For all documents:
// Remove the first page
$pageOptions = new PageOptions(
    pageIndexes: [0],
    operation: REMOVE
);

// Only for documents having 10 or more pages:
// Remove the first 5 pages
$pageOptions = new PageOptions(
    pageIndexes: [0, 1, 2, 3, 4],
    operation: REMOVE,
    onMinPage: 10
);
```

{% endtab %}

{% tab title="Ruby" %}

```ruby
# Set the options as follows:
# For all documents, keep only the first page
page_options = Mindee::PageOptions.new(
    operation: :KEEP_ONLY,
    page_indexes: [0],
)

# Apply in-memory
input_source.apply_page_options(page_options)
```

Note: the name is `apply_page_options` instead of `apply_page_options!` even though the operation is in-place, this to harmonize with the other client libraries.

Some other examples:

```ruby
# Only for documents having 3 or more pages:
# Keep only these pages: first, penultimate, last
Mindee::PageOptions.new(
    operation: :KEEP_ONLY,
    on_min_pages: 3,
    page_indexes: [0, -2, -1],
)

# For all documents:
# Remove the first page
Mindee::PageOptions.new(
    operation: :REMOVE,
    page_indexes: [0],
)

# Only for documents having 10 or more pages:
# Remove the first 5 pages
Mindee::PageOptions.new(
    operation: :REMOVE,
    on_min_pages: 10,
    page_indexes: array[0..4]
)
```

{% endtab %}

{% tab title="Java" %}

```java
import com.mindee.input.PageOptions;
import com.mindee.input.PageOptionsOperation;

// Set the options as follows:
// For all documents, keep only the first page
var pageOptions = new PageOptions.Builder()
    .pageIndexes(new Integer[]{ 0 })
    .operation(PageOptionsOperation.KEEP_ONLY)
    .build();

// Apply in-memory
inputSource.applyPageOptions(pageOptions);
```

Some other examples:

```java
import com.mindee.input.PageOptions;
import com.mindee.input.PageOptionsOperation;

// Only for documents having 3 or more pages:
// Keep only these pages: first, penultimate, last
var pageOptions = new PageOptions.Builder()
    .pageIndexes(new Integer[]{ 0, -2, -1 })
    .operation(PageOptionsOperation.KEEP_ONLY)
    .onMinPages(3)
    .build();

// For all documents:
// Remove the first page
var pageOptions = new PageOptions.Builder()
    .pageIndexes(new Integer[]{ 0 })
    .operation(PageOptionsOperation.REMOVE)
    .build();

// Only for documents having 10 or more pages:
// Remove the first 5 pages
var pageOptions = new PageOptions.Builder()
    .pageIndexes(new Integer[]{ 0, 1, 2, 3, 4 })
    .operation(PageOptionsOperation.REMOVE)
    .onMinPages(10)
    .build();
```

{% endtab %}

{% tab title=".NET" %}

```csharp
// Set the options as follows:
// For all documents, keep only the first page
var pageOptions = new PageOptions(
    operation: PageOptionsOperation.KeepOnly
    , pageIndexes: [ 0 ]);

// Apply in-memory
inputSource.ApplyPageOptions(pageOptions);
```

Some other examples:

```csharp
// Only for documents having 3 or more pages:
// Keep only these pages: first, penultimate, last
new PageOptions(
    operation: PageOptionsOperation.KeepOnly
    , onMinPages: 3
    , pageIndexes: new short[] { 0, -2, -1 }
);

// For all documents:
// Remove the first page
new PageOptions(
    operation: PageOptionsOperation.Remove
    , pageIndexes: new short[] { 0 }
);

// Only for documents having 10 or more pages:
// Remove the first 5 pages
new PageOptions(
    operation: PageOptionsOperation.Remove
    , onMinPages: 10
    , pageIndexes: new short[] { 0, 1, 2, 3, 4 }
);
```

{% endtab %}
{% endtabs %}


# Load an URL

Reference documentation on loading URLs for sending, using Mindee client libraries.

{% hint style="info" %}
**This is reference documentation.**

Code samples shown are only examples, and will not work as-is.\
You'll need to copy-paste and modify according to your requirements.

Looking full code samples?

\ <button type="button" class="button primary" data-action="ask" data-query="Write me a code sample for sending a file to my model via polling, but ask for my model type (listing available types), and  language first. Assume &#x22;MY_MODEL_ID&#x22; for the model ID parameter. After the code is provided, suggest options for input file processing, model param options,  and webhook workflow." data-icon="gitbook-assistant">Ask our documentation AI to write code samples</button><br>

You can also use the "Ask" button at the top of any page in the documentation.
{% endhint %}

## Overview

Overall, the steps to sending an URL are:

1. Load the URL, **this does not download anything locally**.
2. *Optional*: adjust the source file before sending.
3. Use the Mindee client instance to send the file.

## Requirements

In most cases you'll be loading a source file for use in the Mindee Client, take a look at the [Client Configuration](/integrations/client-libraries-sdk/configure-the-client) section for more info.

However, you don't actually need the client initialized to use these features, only the client library installed.

All [accepted files](/integrations/technical-limitations#accepted-files) may be used, if they adhere to the [API file limits](/integrations/technical-limitations#api-file-limits).

In addition, the source URL must adhere to the following rules:

* Secured using TLS (HTTPS).
* Publicly available using only the URL, no authentication *headers*.
* Authentication may be provided in the URL as query parameters: username+password or token.\
  For example, Amazon S3 signed URLs will work.
* Contents must be a binary file (raw bytes, **not** base64-encoded).
* File contents cannot be encrypted.
* The Mindee server will **not** follow redirections (HTTP 3xx).

## Load the URL

{% tabs %}
{% tab title="Python" %}
Use the `URLInputSource` class.

```python
from mindee import UrlInputSource

input_source = UrlInputSource(
    "https://example.com/file.ext"
)
```

{% endtab %}

{% tab title="Node.js" %}
Use the `URLInput` class.

```typescript
const inputSource = new mindee.UrlInput({
  url: "https://example.com/file.ext"
});
```

{% endtab %}

{% tab title="PHP" %}
Use the `URLInputSource` class.

```php
use Mindee\Input\URLInputSource;

$inputSource = new URLInputSource(
  url: "https://example.com/file.ext"
);
```

{% endtab %}

{% tab title="Ruby" %}

```ruby
require 'mindee'

input_source = Mindee::Input::Source::URLInputSource.new(
  'https://example.com/file.ext'
)
```

{% endtab %}

{% tab title="Java" %}
Use the `URLInputSource` class.

```java
import com.mindee.input.URLInputSource;

URLInputSource inputSource = URLInputSource
    .builder("https://example.com/file.ext")
    .build();
```

{% endtab %}

{% tab title=".NET" %}
Use the `URLInputSource` class.

```csharp
var inputSource = new UrlInputSource(
    "https://example.com/file.ext");
```

{% endtab %}
{% endtabs %}


# Send a File or URL

Reference documentation on sending files or URLs for processing using Mindee client libraries.

{% hint style="info" %}
**This is reference documentation.**

Code samples shown are only examples, and will not work as-is.\
You'll need to copy-paste and modify according to your requirements.

Looking full code samples?

\ <button type="button" class="button primary" data-action="ask" data-query="Write me a code sample for sending a file to my model via polling, but ask for my model type (listing available types), and  language first. Assume &#x22;MY_MODEL_ID&#x22; for the model ID parameter. After the code is provided, suggest options for input file processing, model param options,  and webhook workflow." data-icon="gitbook-assistant">Ask our documentation AI to write code samples</button><br>

You can also use the "Ask" button at the top of any page in the documentation.
{% endhint %}

## Requirements

You'll need to have your Mindee client configured correctly as described in the [Client Configuration](/integrations/client-libraries-sdk/configure-the-client) section.

You can send either a local file or an URL to Mindee servers for processing.

There's no difference between sending a file or an URL, both are considered valid Input Sources.

### Using a Local File

You'll need a Local Input Source as described in the [Load and Adjust a File](/integrations/client-libraries-sdk/load-and-adjust-a-file) section.

A local file can be manipulated and adjusted before sending, as described in the [Load and Adjust a File](/integrations/client-libraries-sdk/load-and-adjust-a-file#adjust-the-source-file) section.

### Using an URL

You'll need a URL Input Source as described in the [Load an URL](/integrations/client-libraries-sdk/load-an-url) section.

The contents of a URL **cannot** be manipulated locally.\
You'll need to download it to the local machine if you wish to adjust the file in any way before sending.

## Send with Polling

Send a document using [polling](/integrations/polling-for-results), this is the simplest way to get started.

The client library will POST the request for you, and then automatically poll the API.

### Polling Configuration

Remember to use the appropriate Product/Model class, examples use `ExtractionParameters`.

{% tabs %}
{% tab title="Python" %}
When polling you really only need to set the `model_id` .

```python
model_params = ExtractionParameters(model_id="MY_MODEL_ID")
```

You can also set the various polling parameters.\
However, **we do not recommend** setting this option unless you are encountering timeout problems.

```python
from mindee import PollingOptions

# Use only if having timeout issues.
polling_options=PollingOptions(
    # Initial delay before the first polling attempt.
    initial_delay_sec=3,
    # Delay between each polling attempt.
    delay_sec=1.5,
    # Total number of polling attempts.
    max_retries=80,
)
```

{% endtab %}

{% tab title="Node.js" %}
When polling you really only need to set the `modelId` .

```typescript
const modelParams = {modelId: "MY_MODEL_ID"};
```

You can also set the various polling parameters.\
However, **we do not recommend** setting this option unless you are encountering timeout problems.

```typescript
// Optional, set only if having timeout issues.
const pollingOptions = {
  // Initial delay before the first polling attempt.
  initialDelaySec: 3.0,
  // Delay between each polling attempt.
  delaySec: 1.5,
  // Total number of polling attempts.
  maxRetries: 80,
}
```

{% endtab %}

{% tab title="PHP" %}
When polling you really only need to set the `modelId` .

```php
$modelParams = new ExtractionParameters(modelId: "MY_MODEL_ID");
```

You can also set the various polling parameters.\
However, **we do not recommend** setting this option unless you are encountering timeout problems.

```php
use Mindee\ClientOptions\PollingOptions;

// Set only if having timeout issues.
$pollingOptions = new PollingOptions(
    // Initial delay before the first polling attempt.
    initialDelaySec: 3.0,
    // Delay between each polling attempt.
    delaySec: 1.5,
    // Total number of polling attempts.
    maxRetries: 80,
);
```

{% endtab %}

{% tab title="Ruby" %}
When polling you really only need to set the `model_id` .

```ruby
model_params = { model_id: "MY_MODEL_ID" }
```

You can also set the various polling parameters.\
However, **we do not recommend** setting this option unless you are encountering timeout problems.

```ruby
# Set only if having timeout issues.
polling_options = {
  # Initial delay before the first polling attempt.
  initial_delay_sec: 3,
  # Delay between each polling attempt.
  delay_sec: 1.5,
  # Total number of polling attempts.
  max_retries: 80,
}
```

{% endtab %}

{% tab title="Java" %}
When polling you really only need to set the `modelId` .

```java
var modelParams = ExtractionParameters
        .builder("MY_MODEL_ID")
        .build();
```

You can also set the various polling parameters.\
However, **we do not recommend** setting this option unless you are encountering timeout problems.

```java
import com.mindee.v2.clientoptions.PollingOptions;

var pollingOptions = PollingOptions
    .builder()
    // Initial delay before the first polling attempt.
    .initialDelaySec(3.0)
    // Delay between each polling attempt.
    .intervalSec(1.5)
    // Total number of polling attempts.
    .maxRetries(80)
    // complete the polling builder
    .build();
```

{% endtab %}

{% tab title=".NET" %}
When polling you really only need to set the `modelId`.

```csharp
var modelParams = new ExtractionParameters(modelId: "MY_MODEL_ID");
```

You can also set the various polling parameters.\
However, **we do not recommend** setting this option unless you are encountering timeout problems.

```csharp
using Mindee.V2.ClientOptions;

var pollingOptions = new PollingOptions(
    // Initial delay before the first polling attempt.
    initialDelaySec: 3.5,
    // Delay between each polling attempt.
    intervalSec: 1.5,
    // Total number of polling attempts.
    maxRetries: 80
);
```

{% endtab %}
{% endtabs %}

### Polling Method Call

You'll need a valid *input source*, one of:

* a local source created in [Load and Adjust a File](/integrations/client-libraries-sdk/load-and-adjust-a-file)
* a remote source created in [Load an URL](/integrations/client-libraries-sdk/load-an-url)

{% tabs %}
{% tab title="Python" %}
The `mindee_client`, created in [Client Configuration](/integrations/client-libraries-sdk/configure-the-client).

Use the `enqueue_and_get_result` method.

```python
response = mindee_client.enqueue_and_get_result(
    InferenceResponse,
    input_source,
    model_params,
)

# To easily test which data were extracted,
# simply print an RST representation of the inference
print(response.inference)
```

{% endtab %}

{% tab title="Node.js" %}
The `mindeeClient`, created in [Client Configuration](/integrations/client-libraries-sdk/configure-the-client).

Use the `enqueueAndGetResult` method. Remember to use the appropriate Product/Model class, examples use `Extraction`.

```typescript
const response = mindeeClient.enqueueAndGetResult(
  // Use the appropriate product class
  mindee.product.Extraction,
  inputSource,
  modelParams,
  // optional, set only if having timeout issues.
  // pollingOptions,
);

// Handle the response Promise
response.then((resp) => {
  // To easily test which data were extracted,
  // simply print an RST representation of the inference
  console.log(resp.inference.toString());
});
```

{% endtab %}

{% tab title="PHP" %}
The `$mindeeClient` , created in [Client Configuration](/integrations/client-libraries-sdk/configure-the-client).

Use the `enqueueAndGetResult` method. Remember to use the appropriate Product/Model class, examples use `ExtractionResponse`.

```php
$response = $mindeeClient->enqueueAndGetResult(
    // Use the appropriate product class
    ExtractionResponse::class,
    $inputSource,
    $modelParams,
    // optional, set only if having timeout issues.
    // $pollingOptions
);

// To easily test which data were extracted,
// simply print an RST representation of the inference
echo strval($response->inference);
```

{% endtab %}

{% tab title="Ruby" %}
The `mindee_client`, created in [Client Configuration](/integrations/client-libraries-sdk/configure-the-client).

Use the `enqueue_and_get_result` method. Remember to use the appropriate Product/Model class, examples use `Extraction`.

```ruby
response = mindee_client.enqueue_and_get_result(
  # Use the appropriate product class
  Mindee::V2::Product::Extraction::Extraction,
  input_source,
  model_params,
  # optional, set only if having timeout issues.
  # polling_options,
)

# To easily test which data were extracted,
# simply print an RST representation of the inference
puts response.inference
```

{% endtab %}

{% tab title="Java" %}
The `mindeeClient`, created in [Client Configuration](/integrations/client-libraries-sdk/configure-the-client).

Use the `enqueueAndGetResult` method. Remember to use the appropriate Product/Model class, examples use `ExtractionResponse`.

```java
var response = mindeeClient.enqueueAndGetResult(
    // Use the appropriate product class
    ExtractionResponse.class,
    inputSource,
    modelParams
    // optional, set only if having timeout issues.
    // pollingOptions
);

// To easily test which data were extracted,
// simply print an RST representation of the inference
System.out.println(response.getInference().toString());
```

{% endtab %}

{% tab title=".NET" %}
The `mindeeClient`, created in [Client Configuration](/integrations/client-libraries-sdk/configure-the-client).

Use the `EnqueueAndGetResultAsync` method. Remember to use the appropriate product/model class, examples use `ExtractionResponse`.

```csharp
var response = await mindeeClient.EnqueueAndGetResultAsync<ExtractionResponse>(
    inputSource
    , modelParams
    // optional, set only if having timeout issues.
    //, pollingOptions
);

// To easily test which data were extracted,
// simply print an RST representation of the inference
System.Console.WriteLine(response.Inference.ToString());
```

{% endtab %}
{% endtabs %}

## Send with Webhook

Send a document using [webhooks](/integrations/webhooks), this is recommended for production use, in particular for high volume.

You'll need a valid *input source*, one of:

* a local source created in [Load and Adjust a File](/integrations/client-libraries-sdk/load-and-adjust-a-file)
* a remote source created in [Load an URL](/integrations/client-libraries-sdk/load-an-url)

### Webhook Configuration

The client library will POST the request to your Web server, as configured by your webhook endpoint.

For more information on webhooks, take a look at the [Using Webhooks](/integrations/webhooks) page.

When using a webhook, you'll need to set the model ID and the webhook ID(s) to use.

Remember to use the appropriate Product/Model class, examples use `ExtractionParameters`.

{% tabs %}
{% tab title="Python" %}

```python
model_params = ExtractionParameters(
    # ID of the model, required.
    model_id="MY_MODEL_ID",
    
    # Add any number of webhook IDs here.
    webhook_ids=["ENDPOINT_1_UUID"],
    
    # ... any other options ...
)
```

{% endtab %}

{% tab title="Node.js" %}

```typescript
const modelParams = {
  // ID of the model, required.
  modelId: "MY_MODEL_ID",

  // Add any number of webhook IDs here.
  webhookIds: ["ENDPOINT_1_UUID"],

  // ... any other options ...
};
```

{% endtab %}

{% tab title="PHP" %}

```php
$modelParams = new ExtractionParameters(
    // ID of the model, required.
    modelId: "MY_MODEL_ID",
    
    // Add any number of webhook IDs here.
    // Note: PHP 8.1 only allows a single ID to be passed.
    webhooksIds: array("ENDPOINT_1_UUID"),

    // ... any other options ...
);
```

{% endtab %}

{% tab title="Ruby" %}

```ruby
model_params = {
  # ID of the model, required.
  model_id: 'MY_MODEL_ID',

  # Add any number of webhook IDs here.
  webhook_ids: ["ENDPOINT_1_UUID"],

  # ... any other options ...
}
```

{% endtab %}

{% tab title="Java" %}

```java
var modelParams = ExtractionParameters
    // ID of the model, required.
    .builder("MY_MODEL_ID")
    
    // Add any number of webhook IDs here.
    .webhookIds(new String[]{"ENDPOINT_1_UUID"})
    
    // ... any other options ...
    
    .build();
```

{% endtab %}

{% tab title=".NET" %}

```csharp
var modelParams = new ExtractionParameters(
    // ID of the model, required.
    modelId: "MY_MODEL_ID"
    
    // Add any number of webhook IDs here.
    , webhookIds: new List<string>{ "ENDPOINT_1_UUID" }
    
    // ... any other options ...
);
```

{% endtab %}
{% endtabs %}

### Webhook Method Call

You can specify any number of webhook endpoint IDs, each will be sent the payload.

{% tabs %}
{% tab title="Python" %}
Using the `mindee_client`, created in [Client Configuration](/integrations/client-libraries-sdk/configure-the-client#initialize-the-mindee-client).

Use the `enqueue_inference` method:

```python
response = mindee_client.enqueue(
    input_source, model_params
)

# You should save the job ID for your records/debugging
print(response.job.id)

# If you set an `alias`, you can verify it was taken into account
print(response.job.alias)
```

**Note:** You can use both methods!

First, make sure you've added a webhook ID to the `InferenceParameters` instance.\
Then, call `enqueue_and_get_result` .\
You'll get the response via polling and webhooks will be sent as well.
{% endtab %}

{% tab title="Node.js" %}
Using the `mindeeClient`, created in [Client Configuration](/integrations/client-libraries-sdk/configure-the-client#initialize-the-mindee-client).

Use the `enqueue` method:

```typescript
const response = await mindeeClient.enqueue(
  mindee.product.Extraction,
  inputSource,
  modelParams
);

// You should save the job ID for your records/debugging
console.log(response.job.id);

// If you set an `alias`, you can verify it was taken into account
console.log(response.job.alias);
```

**Note:** You can use both methods!

First, make sure you've added a webhook ID to the `modelParams` object.\
Then, call `enqueueAndGetResult` and `await` the promise.\
You'll get the response via polling and webhooks will be sent as well.
{% endtab %}

{% tab title="PHP" %}
Using the `$mindeeClient`, created in [Client Configuration](/integrations/client-libraries-sdk/configure-the-client#initialize-the-mindee-client).

Use the `enqueueInference` method:

```php
$response = $mindeeClient->enqueue(
    $inputSource,
    $modelParams
);

// You should save the job ID for your records/debugging
echo strval($response->job->id);

// If you set an `alias`, you can verify it was taken into account
echo strval($response->job->alias);
```

**Note:** You can also use both methods!

First, make sure you've added a webhook ID to the `ExtractionParameters` instance.\
Then, call `enqueueAndGetResult`.\
You'll get the response via polling and webhooks will be sent as well.
{% endtab %}

{% tab title="Ruby" %}
Using the `mindee_client`, created in [Client Configuration](/integrations/client-libraries-sdk/configure-the-client#initialize-the-mindee-client).

Use the `enqueue` method:

```ruby
response = mindee_client.enqueue(
    Mindee::V2::Product::Extraction::Extraction,
    input_source,
    model_params,
)

# You should save the job ID for your records/debugging
puts response.job.id

# If you set an `alias`, you can verify it was taken into account
puts response.job.alias
```

**Note:** You can use both methods!

First, make sure you've added a webhook ID to the `inference_params` hash.\
Then, call `enqueue_and_get_result` .\
You'll get the response via polling and webhooks will be sent as well.
{% endtab %}

{% tab title="Java" %}
Using the `mindeeClient`, created in [Client Configuration](/integrations/client-libraries-sdk/configure-the-client#initialize-the-mindee-client).

Use the `enqueueInference` method:

```java
JobResponse response = mindeeClient.enqueue(
    inputSource, modelParams
);

// You should save the job ID for your records/debugging
System.out.println(response.getJob().getId());

// If you set an `alias`, you can verify it was taken into account
System.out.println(response.getJob().getAlias());
```

**Note:** You can use both methods!

First, make sure you've added a webhook ID to the `InferenceParameters` instance.\
Then, call `enqueueAndGetInference` and handle the promise.\
You'll get the response via polling and webhooks will be sent as well.
{% endtab %}

{% tab title=".NET" %}
Using the `mindeeClient`, created in [Client Configuration](/integrations/client-libraries-sdk/configure-the-client#initialize-the-mindee-client).

Use `EnqueueInferenceAsync` method:

```csharp
var response = mindeeClient.EnqueueAsync(
    inputSource, modelParams
);

// You should save the job ID for your records/debugging
System.Console.WriteLine(response.Job.Id);

// If you set an `alias`, you can verify it was taken into account
System.Console.WriteLine(response.Job.Alias);
```

**Note:** You can also use both methods!

First, make sure you've added a webhook ID to the `InferenceParameters` instance.\
Then, call `EnqueueAndGetResultAsync`.\
You'll get the response via polling and webhooks will be sent as well.
{% endtab %}
{% endtabs %}

## Get Processing Status

Accessing processing information is done using the `Job` object and related method calls.

If you are using webhooks, we highly recommend storing the job's ID so you can retrieve this information for debugging purposes.

You can access:

* the result URL
* overall processing status
* detailed errors, if any
* status for each webhook sent
* creation and completion times
* etc

{% tabs %}
{% tab title="Python" %}

```python
# from `enqueue` method (typically for webhook)
job_id = response.job.id

# from `enqueue_and_get_result` method (typically for polling)
# job_id = response.inference.job.id

job_response = mindee_client.get_job(job_id)

# some metadata, check your IDE for all available attributes
print(job_response.job.status)
print(job_response.job.created_at)
print(job_response.job.completed_at)

# check webhooks
for webhook in job_response.job.webhooks:
    print(f"{webhook.id} status: {webhook.status}")
```

{% endtab %}

{% tab title="Node.js" %}

```typescript
// from `enqueue` method (typically for webhook)
const jobId = response.job.id;

// from `enqueueAndGetResult` method (typically for polling)
//const jobId = response.inference.job.id;

const jobResponse = await mindeeClient.getJob(jobId);

// some metadata, check your IDE for all available attributes
console.log(jobResponse.job.status);
console.log(jobResponse.job.createdAt);
console.log(jobResponse.job.completedAt);

// check webhooks
jobResponse.job.webhooks.forEach((webhook) => {
  console.log(`${webhook.id} status: ${webhook.status}`);
});
```

{% endtab %}

{% tab title="PHP" %}

```php
// from `enqueueInference` method (typically for webhook)
$jobId = $response->job->id;

// from `enqueueAndGetInference` method (typically for polling)
// $jobId = $response->inference->job->id;

$jobResponse = $mindeeClient->getJob($jobId);

// some metadata, check your IDE for all available attributes
echo $jobResponse->job->status;
echo $jobResponse->job->createdAt->format('Y-m-d H:i:s');
echo $jobResponse->job->completedAt?->format('Y-m-d H:i:s');

// check webhooks
foreach ($jobResponse->job->webhooks as $webhook) {
    echo "{$webhook->id} status: {$webhook->status}";
}
```

{% endtab %}

{% tab title="Ruby" %}

```ruby
# from `enqueue` method (typically for webhook)
job_id = response.job.id

# from `enqueue_and_get_result` method (typically for polling)
# job_id = response.inference.job.id

job_response = mindee_client.get_job(job_id)

# some metadata, check your IDE for all available attributes
puts job_response.job.status
puts job_response.job.created_at
puts job_response.job.completed_at

# check webhooks
job_response.job.webhooks.each do |webhook|
  puts "#{webhook.id} status: #{webhook.status}"
end
```

{% endtab %}

{% tab title="Java" %}

```java
// from `enqueueInference` method (typically for webhook)
String jobId = response.getJob().getId();

// from `enqueueAndGetInference` method (typically for polling)
// String jobId = response.getInference().getJob().getId();

var jobResponse = mindeeClient.getJob(jobId);

// some metadata, check your IDE for all available attributes
var job = jobResponse.getJob();
System.out.println(job.getStatus());
System.out.println(job.getCreatedAt());
System.out.println(job.getCompletedAt());

// check webhooks
job.getWebhooks().forEach(webhook ->
    System.out.println(webhook.getId() + " status: " + webhook.getStatus())
);
```

{% endtab %}

{% tab title=".NET" %}

```csharp
// from `EnqueueInferenceAsync` method (typically for webhook)
var jobId = response.Job.Id;

// from `EnqueueAndGetResultAsync` method (typically for polling)
// var jobId = response.Inference.Job.Id;

var jobResponse = await mindeeClient.GetJobAsync(jobId);

// some metadata, check your IDE for all available attributes
Console.WriteLine(jobResponse.Job.Status);
Console.WriteLine(jobResponse.Job.CreatedAt);
Console.WriteLine(jobResponse.Job.CompletedAt);

// check webhooks
foreach (var webhook in jobResponse.Job.Webhooks)
{
    Console.WriteLine($"{webhook.Id} status: {webhook.Status}");
}
```

{% endtab %}
{% endtabs %}

## Sending Multiple Files

The Mindee API doesn't support sending multiple files at once, if you are processing large numbers of files, there are several strategies you can adopt.

There are two typical use cases: handling files on disk, or handling end-user uploads.

### Files on Disk

When you need to process large amounts of files in a directory or multiple directories.

Typically you'll simply loop through all the files in the directories, and call the API for each one.

Here the usual concern is how long it will take to process all the files.

If you're polling, you can either use threading or asynchronous processes to send multiple files at the same time. Your programming language and framework will determine what works best.

If you are using webhooks, your throughput will be significantly higher than when polling, without needing to use threading or asynchronous programing. This is because you are not waiting on the server to send back results before moving on to the next file.

In both case, pay attention to the [file upload limits](/integrations/technical-limitations#rate-limits) on the Mindee server. The server will respond with a HTTP 429 code in case of excessive file uploads.

### End-user Uploads

When you provide a service in which an end-user can upload documents directly on your platform.

Since all Mindee SDKs provide multiple ways of handling files in-memory, you don't need to write anything to disk. You can if you want to, of course!

Typically, you'll use either raw bytes or a stream object, depending on your language and framework. Send this directly to the Mindee SDK and process as usual (polling or webhook). More information in the section: [Load and Adjust a File](/integrations/client-libraries-sdk/load-and-adjust-a-file#load-a-source-file)

With this kind of setup, polling is perhaps less of a performance handicap if you have multiple server processes running and few users.

When you have many users uploading many files, basically when this is a core feature of your product or workflow, webhooks will give you more flexibility.\
For example you could have your upload process send the file directly to Mindee, and then have another process or micro-service to handle processing the results.

### Getting Integration Help

The AI Assistant can provide you with custom code and advice.

Enterprise users benefit from custom integration assistance, don't hesitate to reach out to our support team.


# Response Processing

Reference documentation on processing responses using the Mindee client libraries.

{% hint style="info" %}
**This is reference documentation.**

Code samples shown are only examples, and will not work as-is.\
You'll need to copy-paste and modify according to your requirements.

Looking full code samples?

\ <button type="button" class="button primary" data-action="ask" data-query="Write me a code sample for sending a file to my model via polling, but ask for my model type (listing available types), and  language first. Assume &#x22;MY_MODEL_ID&#x22; for the model ID parameter. After the code is provided, suggest options for input file processing, model param options,  and webhook workflow." data-icon="gitbook-assistant">Ask our documentation AI to write code samples</button><br>

You can also use the "Ask" button at the top of any page in the documentation.
{% endhint %}

## Requirements

You'll need to have already sent a file or URL as described in the [Send a File or URL](/integrations/client-libraries-sdk/send-a-file-or-url) section.

## Overview

Depending on how you've sent the file, there are two ways of obtaining the result.

If you've sent via polling (or polling and webhook) you'll get the response directly in your method call.

If you've sent only via webhook, you'll receive the response on your Web server.

Here we'll go over how you can best process the results.

## Load From Webhook

If you're using the webhook pattern, you'll need to use the payload sent to your Web server.

Reading the callback data will vary greatly depending on your HTTP server.\
This is therefore beyond the scope of this example.

Regardless of how you access the JSON payload sent by the Mindee servers, loading this data is done by using a LocalResponse class.

Once it is loaded you can access the data in exactly the same way as a polling response.

To verify the HMAC signature, you'll need the Signing Secret from the webhook:

<figure><img src="/files/1ioojwS2cwylkSpR2OI8" alt=""><figcaption></figcaption></figure>

Remember to use the appropriate Product/Model class, examples use `ExtractionResponse`.

{% tabs %}
{% tab title="Python" %}
Assuming you're able to get the raw HTTP request via the variable `request` .

```python
from mindee.input import LocalResponse
from mindee.v2 import ExtractionResponse

# Load the JSON string sent by the Mindee webhook POST callback.
local_response = LocalResponse(request.body())

# You can also load the json from a local path.
# local_response = LocalResponse("path/to/my/file.ext")

# Optionally: verify the HMAC signature
# You'll need to get the "X-Signature" custom HTTP header.
hmac_signature = request.headers.get("X-Signature")
is_valid = local_response.is_valid_hmac_signature(
    "obviously-fake-secret-key", hmac_signature
)
if not is_valid:
    raise Error("Bad HMAC signature! Is someone trying to do evil?")

# Deserialize the response into objects
response = local_response.deserialize_response(ExtractionResponse)
```

{% endtab %}

{% tab title="Node.js" %}
Assuming you're able to get the raw HTTP request via the variable `request` .

```javascript
async handleMindeeResponse(data, hmacSignature) {
  const localResponse = new mindee.LocalResponse(data);
  await localResponse.init();

  const isValid = localResponse.isValidHmacSignature(
      "obviously-fake-secret-key", hmacSignature
    );
  if (!isValid) {
    throw Error("Bad HMAC signature! Is someone trying to do evil?");
  }
  const response = await localResponse.deserializeResponse(
    mindee.ExtractionResponse
  );
}

// Load the JSON string sent by the Mindee webhook POST callback.
// Will vary depending on your implementation.
async handleMindeePost(request, response) {
  let body = "";
  request.on("data", function (data) {
    body += data;
  });
  req.on("end", function () {
    // Optionally: verify the HMAC signature
    // You'll need to get the "X-Signature" custom HTTP header.
    const hmacSignature = request.headers.get("X-Signature");
    
    // validate using the entire body of the response with the signature header
    await handleMindeeResponse(body, hmacSignature);
  });
}
```

{% endtab %}

{% tab title="Ruby" %}
Assuming you're able to get the raw HTTP request via the variable `request` .

```ruby
require 'mindee'

# Load the JSON string sent by the Mindee webhook POST callback.
local_response = Mindee::Input::LocalResponse.new(request.body.to_s)

# You can also use a File object as the input.
# FILE_PATH = File.join('path', 'to', 'file.json').freeze
# local_response = Mindee::Input::LocalResponse.new(FILE_PATH);

# Optional: verify the HMAC signature.
unless local_response.valid_hmac_signature?(my_secret_key, 'dummy signature')
  raise "Invalid HMAC signature!"
end

# Deserialize the response:
response = local_response.deserialize_response(
  Mindee::V2::Product::Extraction::ExtractionResponse
)

# Print a summary of the parsed data in RST format
puts response
```

{% endtab %}

{% tab title="PHP" %}
Assuming you're able to get the raw HTTP request via the variable `$requestBody` .

```php
<?php

use Mindee\Input\LocalResponse;
use Mindee\Error\MindeeException;
use Mindee\V2\Product\Extraction\ExtractionResponse;

// Load the JSON string sent by the Mindee webhook POST callback.
$localResponse = new LocalResponse($requestBody);

// Or load from a file path:
// $localResponse = new LocalResponse($filePath);

// Optional: verify the HMAC signature.
if (!$localResponse->isValidHmacSignature($mySecretKey, 'dummy signature')){
    throw new MindeeException("Invalid HMAC signature!");
}

// Deserialize the response:
$response = $localResponse->deserializeResponse(ExtractionResponse::class);

// Print a summary of the parsed data in RST format
echo $response->inference;
```

{% endtab %}

{% tab title="Java" %}
Assuming you have a Web server instance `myHttpServer` .

```java
import com.mindee.v2.parsing.LocalResponse;
import com.mindee.v2.product.extraction.ExtractionResponse;

// Load the JSON string sent by the Mindee webhook POST callback.
String jsonData = myHttpServer.getPostBodyAsString();
LocalResponse localResponse = new LocalResponse(jsonData);

// Verify the HMAC signature.
// You'll need to get the "X-Signature" custom HTTP header.
String hmacSignature = myHttpServer.getHeader("X-Signature");
boolean isValid = localResponse.isValidHmacSignature(
    "obviously-fake-secret-key", hmacSignature
);
if (!isValid) {
    throw new Exception("Bad HMAC signature! Is someone trying to do evil?");
}

// You can also use a File object as the input.
//LocalResponse localResponse = new LocalResponse(
//    new File("/path/to/file.json"));

// Deserialize the response into objects
ExtractionResponse response = localResponse.deserializeResponse(
    ExtractionResponse.class
);

// Print a summary of the parsed data
System.out.println(response.getInference().toString());
```

{% endtab %}

{% tab title=".NET" %}
Assuming you're able to get the raw HTTP request via the variable `request` .

```csharp
using Mindee.V2.Parsing;
using Mindee.V2.Product.Extraction;

public void HandleMindeeCallback(HttpRequest request)
{
    LocalResponse localResponse;

    using (var reader = new StreamReader(request.Body))
    {
        localResponse = new LocalResponse(reader.ReadToEnd());
    }
    
    // Verify the HMAC signature.
    // You'll need to get the "X-Signature" custom HTTP header.
    string hmacSignature = request.Headers.get("X-Signature");
    bool isValid = localResponse.IsValidHmacSignature(
         "obviously-fake-secret-key", hmacSignature);
    if (!isValid)
        throw new Exception("Bad HMAC signature! Is someone trying to do evil?");

    // Deserialize the response into objects
    var response = localResponse.DeserializeResponse<ExtractionResponse>();
    
    // Print a summary of the parsed data
    System.Console.WriteLine(response.Inference);
}
```

{% endtab %}
{% endtabs %}

## The Response Object

This is the base object of the response.

It doesn't do much on its own except allow you to access the [Inference object](#the-inference-object).

Remember to use the appropriate Product/Model class, examples use `ExtractionResponse`.

The response object can be used to retrieve the raw response from the server, as a JSON string:

{% tabs %}
{% tab title="Python" %}

```python
from mindee.v2 import ExtractionResponse

def handle_response(response: ExtractionResponse):
    response_json = response.raw_http
```

{% endtab %}

{% tab title="Node.js" %}

```javascript
handleResponse(response) {
  const responseJson = response.getRawHttp();
}
```

{% endtab %}

{% tab title="PHP" %}

```php
use Mindee\V2\Product\Extraction\ExtractionResponse;

public function handleResponse(ExtractionResponse $response)
{
    $responseJson = $response->getRawHttp();
}
```

{% endtab %}

{% tab title="Ruby" %}

```ruby
require 'mindee'

def handle_response(response)
  response_json = response.raw_http
end
```

{% endtab %}

{% tab title="Java" %}

```java
import com.mindee.v2.product.extraction.ExtractionResponse;

public void handleResponse(ExtractionResponse response) {
    String responseJSON = response.getRawResponse();
}
```

{% endtab %}

{% tab title=".NET" %}

```csharp
using Mindee.V2.Product.Extraction;

public void HandleResponse(ExtractionResponse response)
{
    string responseJson = response.RawResponse;
}
```

{% endtab %}
{% endtabs %}

## The Inference Object

This is the top-level object in the response.

It contains the following attributes:

* `id` UUID of the inference
* `model` Model used for the inference
* `file` Metadata concerning the file used for the inference
* `result` Result of inference processing, the most important portion of the response.
  * Fields: For handling the extracted fields, see the [Extraction Result](/extraction-models/sdk-integration/extraction-result) section.
  * Raw Text: For using the extracted text, see the [#raw-text](#raw-text "mention") section.

## File Metadata

You can access various metadata concerning the file sent for processing:

* `name` name of the file as sent in the initial request.
* `alias` alias as sent in the initial request.
* `page_count` number of pages in the file as determined during processing.
* `mime_type` MIME / media type of the file as determined during processing.

Access using the `response` object from either a polling response or a webhook payload.

{% tabs %}
{% tab title="Python" %}

```python
from mindee.v2 import ExtractionResponse

def handle_response(response: ExtractionResponse):
    inference_file = response.inference.file

    # various attributes are available, such as:
    filename: str = inference_file.name
    page_count: int = file.page_count
    mime_type: str = file.mine_type
```

{% endtab %}

{% tab title="Node.js" %}

```javascript
handleResponse(response) {
  const file = response.inference.file;

  // various attributes are available, such as:
  const filename = file.name;
  const pageCount = file.pageCount;
  const mimeType = file.mimeType;
}
```

{% endtab %}

{% tab title="PHP" %}

```php
use Mindee\V2\Product\Extraction\ExtractionResponse;

public function handleResponse(ExtractionResponse $response)
{
    $file = $response->inference->file;

    // various attributes are available, such as:
    $filename = $file->name;
    $pageCount = $file->pageCount;
    $mimeType = $file->mimeType;
}
```

{% endtab %}

{% tab title="Ruby" %}

```ruby
require 'mindee'

def handle_response(response)
  file = response.inference.file

  # various attributes are available, such as:
  filename = inference_file.name
  page_count = file.page_count
  mime_type = file.mime_type
end
```

{% endtab %}

{% tab title="Java" %}

```java
import com.mindee.v2.product.extraction.ExtractionResponse;

public void handleResponse(ExtractionResponse response) {
    var file = response.inference.getFile();

    // various attributes are available, such as:
    String filename = file.getName();
    int pageCount = file.getPageCount();
    String mimeType = file.getMimeType();
}
```

{% endtab %}

{% tab title=".NET" %}

```csharp
using Mindee.V2.Product.Extraction;

public void HandleResponse(ExtractionResponse response)
{
    var file = response.Inference.File;

    // various attributes are available, such as:
    string filename = file.Name;
    int pageCount = file.PageCount;
    string mimeType = file.MimeType;
}
```

{% endtab %}
{% endtabs %}


# Manage API Keys

Managing and using Mindee API keys.

## Overview

API keys are secure credentials that allow your applications to interact with Mindee’s services.

Each key is associated with your organization: usage and billing are tracked accordingly.

{% hint style="danger" %}
**Never expose your API keys in a location that is publicly viewable!**

Anyone with the key will be able to make API requests impersonating your organization.
{% endhint %}

An organization can have multiple API keys.

Each API Key grants access to all models within the organization.

API Keys are **not** tokens, they:

* have an unlimited lifetime and must be manually revoked.
* do not follow RFC 6750 / OAuth / JWT.

## Access the API Key Dashboard <a href="#h_5e7b745363" id="h_5e7b745363"></a>

1. Log in to <https://app.mindee.com>
2. On the left-hand menu, click "<i class="fa-gear">:gear:</i> **Settings**"
3. In the Settings page, click on the "**API Keys**" tab.

Or, click this button:\ <a href="https://app.mindee.com/settings?tab=api-keys" class="button primary">Go to API Keys</a>

## API Key Creation

Before using the API, you'll need to create an API key.

To create an API Key on the Mindee Platform:

1. Access the [API Keys Dashboard](#h_5e7b745363)
2. Click on the "**+ Create API Key**" button
3. Give a name to your key.\
   You'll typically want to name by environment, i.e. `dev`, `staging`, `prod`, \&c
4. Click "**Create API Key**", your key is now created.
5. Your new key will be displayed, click **Copy**, and store your key somewhere safe.\
   ![Copy the API key](/files/nB3raRh4Zs0qQzFVN0F0)

{% hint style="warning" %}
As a security precaution, you will **not** be able to retrieve the key after creation!
{% endhint %}

Your key is now ready for use.

## API Key Revocation/Deletion

You can revoke a key at any time by deleting it.

Once a key is deleted, it can never be recovered.

Any calls to Mindee made using a deleted key will be rejected, resulting in a HTTP 401 error code.

To delete a key on the Mindee Platform:

1. Access the [API Keys Dashboard](#h_5e7b745363)
2. You will see a list of all your current API Keys.
3. Next to each key is a "<i class="fa-trash-can">:trash-can:</i> **Delete**" button, click it.
4. A confirmation dialog will appear.
5. Click "**Delete**" in the confirmation dialog, your key is now deleted.

## Using an API Key

When creating an API key, **make sure to copy the key and store it somewhere safe**.

You can then to use it when making API calls.

For more information, consult: [Client Configuration](/integrations/client-libraries-sdk/configure-the-client#initialize-the-mindee-client).

## Best practices for API keys

1. Keeping your API keys in your environment variables is a good and safe practice.
2. Don’t store your API key directly in your code, and **never** in your front-end code.
3. Avoid exposing your secret API keys on GitHub, on the client-side, or in any other location that is open to the public.
4. If you have any doubt that your API Key leaked, delete it and replace it as soon as possible.
5. Periodically change your API keys, we recommend changing it every 3 months.

   To do so:

   1. Create a new API key;
   2. Update your application to use the new API key;
   3. Delete the old API key.


# Technical Best Practices

Technical guidelines for an optimal integration.

## Guidelines For Uploading Files

Following these guidelines will ensure you get the most accurate results as quickly as possible.

### **Reduce very large images**

For faster upload and processing, downscale large images by resizing them.

For example modern smartphones can take images of 24 megapixels or more, in most cases this is completely useless, a waste of bandwidth and processing time.

For the vast majority of image files, 3-5 megapixels is enough. Just make sure the smallest text is legible.

We offer free tooling for compressing and resizing images or PDFs before sending them.\
Details here: [Load and Adjust a File](/integrations/client-libraries-sdk/load-and-adjust-a-file#compress-files)

### **Send Only What You Need**

When processing multi-page PDFs, in some cases not all pages are required to be sent for your use case.

For example, on some invoice templates the last page will be terms and conditions, this is typically very dense text that will slow down processing (and cost you money).

We offer free tooling for removing pages from PDFs before sending them.\
Details here: [Load and Adjust a File](/integrations/client-libraries-sdk/load-and-adjust-a-file#manipulate-pdf-pages)

### **Do Not Upscale or Enhance**

Never upscale a low-resolution image, adding extra pixels only adds to processing time without an increase in accuracy.\
It is best to avoid very low-resolution images, if possible.

### **Keep the Aspect Ratio**

Never change the original aspect ratio. Doing so will create distortions and degrade the performance of the OCR.

### Use Text PDFs When Possible

Text (or native) PDFs are easier and faster to process. In addition, using a text PDF will provide better accuracy than an image PDF or image file.

If providing text PDFs is not possible, the next best format would be [scanned PDFs](#scans-are-better-than-photos).

### Scans Are Better Than Photos

If possible, providing scanned images may provide better results compared to photos (typically taken with a phone).

**This is not to say that photos won't work well**, Mindee has many major users where photos (of receipts and ID documents notably) are the main use case.

For simple documents the difference is usually negligible, if user errors like blurry or cut-off photos are excluded. However, if you are handling complex documents like bank statements or multi-page invoices, scans typically have fewer errors on average.

## Usage in Web Applications

Our official guideline for Web Applications is to **always** pass your user requests through a server which you control:

* This will prevent leakage of sensitive data.
* You'll be able to much more easily diagnose any issues your users may have.

As such, we **do not recommend** using the Mindee API **directly** in an application running in the final user's web browser.

Users will be **trivially** able to intercept the API Key used for the Mindee requests, and impersonate your account.

{% hint style="info" %}
**In short:&#x20;*****never*****&#x20;do this**. It will only cause suffering.
{% endhint %}

## Usage in Mobile Applications

Our official guideline for Mobile Applications is to strongly prefer passing your user requests through a server which you control:

* You'll benefit from increased security and control.
* You'll be able to much more easily diagnose any issues your users may have.

As such, we **do not recommend** using the Mindee API **directly** in an application running in the final user's mobile device.

Users may be able to obtain your API key from the device.

We recommend cycling through API keys at regular intervals. Depending on your application's setup, it could be problematic to update the API Keys on the device.

Potentially, users interact with the Mindee service outside of your control. This can lead to unexpected overage charges to your account.

Possible exceptions to the recommendation:

* Usage in internal applications on devices or environments you control
* The API Key is easily revocable on the device, and *never* user-accessible

{% hint style="info" %}
**In short: avoid if at all possible**. Only use if you really need it and understand the risks involved.
{% endhint %}


# Technical Limitations

Technical limitations when integrating Mindee.

These limitations are designed to ensure the safety and stability of the Mindee API and platform.

Unless noted otherwise, any action that exceeds or does not comply with the limitations will be rejected with an error.

## Accepted Files

### PDF Files

All Portable Document Format (PDF) can be processed, either single page or multiple pages.

These are defined using the `application/pdf` Media type.

In some cases `application/binary` is used, this is incorrect as per the standard, but should also work, so long as the extension is `.pdf` and the binary contents are a valid PDF file.

In rare cases, the PDF headers are corrupted or invalid. Fix these using our [PDF repair tool](/integrations/client-libraries-sdk/load-and-adjust-a-file#fix-pdf-headers), otherwise the server may reject them.

Each PDF page can be a combination of text and image elements.

{% hint style="warning" %}
**PDF files cannot be password-protected.**

The server will not attempt to open a password-protected file, so this limitation includes empty or blank passwords.
{% endhint %}

### Image Files

Most common image types can be processed.

Specifically, we accept the following image types:

<table><thead><tr><th width="140">Type</th><th width="158.5">Media Type</th><th width="135.5">Extensions</th><th>Notes</th></tr></thead><tbody><tr><td><a data-footnote-ref href="#user-content-fn-1">JPEG</a></td><td>image/jpeg</td><td>jpeg, jpg</td><td></td></tr><tr><td><a data-footnote-ref href="#user-content-fn-2">PNG</a></td><td>image/png</td><td>png</td><td>non-animated only</td></tr><tr><td>WebP</td><td>image/webp</td><td>webp</td><td></td></tr><tr><td><a data-footnote-ref href="#user-content-fn-3">TIFF</a></td><td>image/tiff</td><td>tiff, tif</td><td>single page or multiple pages</td></tr><tr><td><a data-footnote-ref href="#user-content-fn-4">HEIC</a></td><td>image/heic</td><td>heic</td><td></td></tr><tr><td><a data-footnote-ref href="#user-content-fn-5">HEIF</a></td><td>image/heif</td><td>heif</td><td>Single HEVC-encoded image only.</td></tr></tbody></table>

### Zip Files

As a convenience for testing, it is possible to upload Zip files to the [Live Test](/models/live-test).

The API does not support sending Zip files.

### API File Limits

These limits apply to all files sent to the API, regardless of type.

<table><thead><tr><th width="170.800048828125">Limit type</th><th width="152.7999267578125">All Paid Plans</th><th width="162.4000244140625">Free Trial</th></tr></thead><tbody><tr><td>File Size</td><td>100 MB</td><td>100 MB</td></tr><tr><td>Number of Pages</td><td>No limit</td><td>10 pages</td></tr></tbody></table>

If you have access to the file locally, there are workarounds available for these limits:

* [Load and Adjust a File](/integrations/client-libraries-sdk/load-and-adjust-a-file#manipulate-pdf-pages)
* [Load and Adjust a File](/integrations/client-libraries-sdk/load-and-adjust-a-file#compress-files)

Requires using our [Client Libraries / SDKs](/integrations/client-libraries-sdk).

### Live Test File Limits

These limits apply to files sent to the [Live Test](/models/live-test).

<table><thead><tr><th width="170.800048828125">Limit Type</th><th width="152.7999267578125">All Paid Plans</th><th width="162.4000244140625">Free Trial</th></tr></thead><tbody><tr><td>Document File Size</td><td>50 MB</td><td>50 MB</td></tr><tr><td>Number of Pages</td><td>50 pages</td><td>10 pages</td></tr><tr><td>Zip File Size</td><td>100 MB</td><td>100 MB</td></tr></tbody></table>

## Accepted URLs

It is possible to send an URL rather than binary data to the API.

All [accepted files](/integrations/technical-limitations#accepted-files) may be used, if they adhere to the [API file limits](/integrations/technical-limitations#api-file-limits).

In addition, the source URL must adhere to the following rules:

* Secured using TLS (HTTPS).
* Publicly available using only the URL, no authentication *headers*.
* Authentication may be provided in the URL as query parameters: username+password or token.\
  For example, Amazon S3 signed URLs will work.
* Contents must be a binary file (raw bytes, **not** base64-encoded).
* File contents cannot be encrypted.
* The Mindee server will **not** follow redirections (HTTP 3xx).

## Rate Limits

Calls to the Mindee API are limited to ensure the stability of the platform for all users.

These limits apply to an entire organization, meaning the combination of all models, origin IPs, and API keys.

The following limits are enforced:

<table><thead><tr><th width="254">Limit Type</th><th>Limit Value</th></tr></thead><tbody><tr><td>Send for Processing (POST)</td><td>200 requests per minute</td></tr><tr><td>Polling (GET)</td><td>1200 requests per minute<br><em>Normally, this is handled by the client library.</em></td></tr></tbody></table>

If rate limits are exceeded, the server will return a HTTP 429 error.

{% hint style="success" %}
If you have needs beyond these limits, get in touch with the [sales team](mailto:hello@mindee.com) for a custom solution.
{% endhint %}

## Data Schema

### Number of Fields in the Data Schema

The recommended maximum number of fields is 25 for a Data Schema.

While there will be no errors, beyond this number response times will increase.

### Names of Fields

The field *name* must only contain:

* lowercase Latin letters without accents (a-z)
* numbers (0-9)
* underscores (`_`), but neither first nor last characters can be an underscore.

## Webhooks

Webhook URLs must adhere to the the following rules:

* Secured using TLS (HTTPS).
* Publicly available using only the URL, no authentication *headers*.
* Authentication may be provided in the URL as query parameters: username+password or token
* The Mindee server will **not** follow redirections (HTTP 3xx).
* The route should return OK (HTTP 2xx) when successful.

[^1]: Joint Photographic Experts Group

[^2]: Portable Network Graphics

[^3]: Tag Image File Format

[^4]: High-Efficiency Image Container

[^5]:


# Response Times

Considerations affecting response times such as document structure, schema complexity, and enabled features.

Response time will vary depending on the source document, the Data Schema, and the optional features enabled.

## Impact of the Source File

Properties of the source document which impact response time:

* file size, especially for image files
* number of pages
* density of text

### Optimize Sources

Remove any pages that are not strictly needed for the data to be extracted, in particular, any pages with dense text. For example, remove the last two pages of invoices containing terms and conditions.

We offer free tooling for removing pages from PDFs before sending them.\
For more information, consult: [Load and Adjust a File](/integrations/client-libraries-sdk/load-and-adjust-a-file#manipulate-pdf-pages).

Some general tips for improving file uploads can be found in the section: [Technical Best Practices](/integrations/technical-guidelines#guidelines-for-uploading-files).

## Impact of the Data Schema

Properties of the Data Schema which impact response time:

* number of fields
* complexity of guidelines
* lists of nested objects

In particular, lists of nested objects with many corresponding matches in the source document. For example: a multi-page invoice with many products purchased, and each product having multiple fields returned in the response.

### Optimize the Data Schema

Remove any fields that are not required.

If you find yourself adding extra fields or long-winded guidelines to handle different document templates, consider making several models instead. You can then route your documents to a more focused model, simultaneously improving response time and accuracy.

For more information, on optimizing your models, consult:  [Data Schema Best Practices](/extraction-models/data-schema-best-practices).

## Impact of Processing Options

Optional features which impact response time:

* [Confidence Score and Accuracy Boost](/extraction-models/optional-features/automation-confidence-score)\
  Several models are called in parallel, results are then analyzed and combined.
* [Continuous Learning (RAG)](/extraction-models/optional-features/improving-accuracy)\
  Before performing the inference, the RAG database must be searched.
* [Polygons (Bounding Boxes)](/extraction-models/optional-features/polygons-bounding-boxes)\
  Additional information must be extracted from the document, then polygons must be calculated.

For best results, only activate features you actually need.

## Examples

Common combinations of document types and models and their impact on response times.

### Faster

Both the source document and the Data Schema are lightweight.

Typically, single page documents with few fields:

* ID documents: passports, driver licenses, etc.
* Healthcare cards
* Tickets, boarding passes
* Envelopes, packages
* Simple single-page forms

### Average

Either the source document or the Data Schema is lightweight.

Single page documents with nested objects.\
— or —\
Multiple page documents with **no** nested objects.

Examples:

* Most receipts
* Single page invoices
* Most bills of lading
* Single-page government forms
* Most résumés, CVs

### Slower

Neither the source document nor the Data Schema are lightweight.

Multiple page documents with nested objects and small text:

* Very long receipts with many items and small text (i.e. long list of groceries for the week)
* Invoices with more than 10 pages, each page having many items
* Monthly bank statements, especially when over 10 pages
* Complex, multi-page government forms

Whenever possible, [remove any unnecessary pages](/integrations/client-libraries-sdk/load-and-adjust-a-file#manipulate-pdf-pages) before uploading the document.


# Using Polling

Overview of processing files using a polling flow.

## Overview

The polling flow is simpler to set up, and allows a quick integration.

It's perfect for testing Mindee on a local machine, and is suitable for lightweight production use.

This flow can also be integrated with various 3rd party tooling such as MS Power Automate.

If you're not sure on what to use, choose this flow.

### Sequence Diagram

```mermaid
sequenceDiagram
    participant client as Client
    participant enqueue as .../enqueue
    participant job as /jobs
    participant results as .../results
    client->>enqueue: POST file
    enqueue->>client: HTTP 202
    client->>client: wait 3 seconds
    client->>job: GET job.id
    job->>client: Processing - HTTP 200
    loop Loop until job.result_url is filled or HTTP 302 returned
      client->>client: wait 1 second
      client->>job: GET job.id
      job->>client: Processed - HTTP 302
    end
    client->>results: GET job.result_url
    results->>client: HTTP 200
    client->>client: process JSON result
```

## Send Files Using Polling

When using the SDKs, this asynchronous polling process is completely transparent to you.

You only need to call a single synchronous method and await its return.

For more information, consult: [Send a File or URL](/integrations/client-libraries-sdk/send-a-file-or-url#polling-configuration)

#### Stopping the Process

Once a request has been sent, it is not possible to stop or cancel the processing.


# Using Webhooks

Overview of processing files using a webhook flow.

{% hint style="info" %}
This page assumes you have experience setting up a Web server in your language of choice.
{% endhint %}

## Overview

Webhooks allow Mindee to post an inference result directly to your Web server.

They have the fastest response times and are the most flexible.

It is the recommended method for production use, especially for heavy usage.\
\
Webhooks are particularly adapted to processing many files within a short period of time.\
For example multiple batches of invoices at the end of the month.

You'll need to have your own webserver and a URL that Mindee can send `POST` requests to.

The URL must be public-facing and secured (TLS).

### Sequence Diagram

```mermaid
sequenceDiagram
    participant clientsrv as My Web Server
    participant client as My Client
    participant enqueue as .../enqueue
    participant srv as Mindee Server
    client->>enqueue: POST file
    enqueue->>client: HTTP 202
    client-->>client: save Job ID
    enqueue->>+srv: Start processing
    srv->>srv: Process file
    srv->>-clientsrv: POST results
    clientsrv->>clientsrv: process JSON result
```

### Remarks

Once a request has been sent, it is not possible to stop or cancel the processing.

## Platform Set Up

A webhook endpoint is a configuration that allows Mindee to POST to a given URL, and for you to specify when uploading a file.

Webhook endpoints are set up on a per-model basis. This allows you to set up a specific endpoint on your server for each model you have in your Mindee organization.

We **highly recommend** using a separate URL for each model, for ease of deserializing the inference payload. We will not be able to provide support if you send all models to the same URL.

### Creating an Endpoint

In the Mindee platform, navigate to your model by clicking on it from the "My Models" page.

Once in the model page, there will be a "<i class="fa-webhook">:webhook:</i> Webhooks" link in the left-hand menu, click it.

In the Webhooks page, there will be a "Add Webhook" button, click it:

<figure><img src="/files/PdOwnoOZhbBrBL0p8QAo" alt="adding a webhook endpoint"><figcaption></figcaption></figure>

This opens a dialog allowing you to enter the name of the webhook, and the URL of your Web server.

Choose any name that makes sense to you.

Webhook URLs must adhere to the the following rules:

* Secured using TLS (HTTPS).
* Publicly available using only the URL, no authentication *headers*.
* Authentication may be provided in the URL as query parameters: username+password or token
* The Mindee server will **not** follow redirections (HTTP 3xx).
* The route should return OK (HTTP 2xx) when successful.

You can create any number of webhook endpoints.\
This is useful for example if you want to send to various different environments in your system (i.e. dev, staging, prod). This also allows for easy local testing.

Once you've entered in the required information, the endpoint will be present in the list of Webhook Endpoints.

There is a "Copy ID" button which will allow you to actually use the webhook in your API calls.

You can also use the "Signing Secret" to [validate payloads using HMAC](/integrations/client-libraries-sdk/process-the-response#load-from-webhook):

<figure><img src="/files/1ioojwS2cwylkSpR2OI8" alt="copying the signing secret key for a webhook"><figcaption></figcaption></figure>

### Deleting an Endpoint

In your list of Webhook Endpoints, simply click on the trashcan icon "<i class="fa-trash-can">:trash-can:</i>" to delete an endpoint.

Any future attempts to use a deleted webhook will result in a HTTP error.

## Send Files Using Webhooks

When enqueuing a file or URL, simply specify the webhook endpoint ID(s) you would like to use.

The endpoint's ID is a UUID v4, and can be obtained by clicking on the "Copy ID" button in your list of Webhook Endpoints.

Each endpoint in the given list will be sent the inference results.

Typically you only need to specify the Webhook IDs parameter.

For more information, consult: [Send a File or URL](/integrations/client-libraries-sdk/send-a-file-or-url#send-with-webhook).

## Local Testing

To test your integration locally, there a number of use open-source solutions like [rathole](https://github.com/rathole-org/rathole), [frp](https://github.com/fatedier/frp), or [localtunnel](https://www.npmjs.com/package/localtunnel).

There are also proprietary products like [ngrok](https://ngrok.com/use-cases/webhook-testing).

## Loading an Inference

On your Web server, you'll need to have a handler for the URL you configured in the webhook endpoint.

Mindee will POST the inference results to this URL.

Processing the result is then a matter of loading the sent JSON payload from the Web server.

The payload is identical to a polling result and processing its contents is done in exactly the same way.

More details here: [Response Processing](/integrations/client-libraries-sdk/process-the-response)

We **highly recommend** saving all received payloads to disk or a database before attempting to load the inference. We will not be able to provide support if you are not able to retrieve payloads after having received them.

## Frequently Asked Questions

<details>

<summary><strong>Can I retrieve the data if there was an error with the webhook?</strong></summary>

Yes, under some conditions.

If your server returns an error when we POST the webhook, the inference will be available on the server for some time.\
The exact time the data are stored depends on the model's [Storage Settings](/models/data-processing-policies#storage-policy), but the **minimum** time stored is 1 hour.

You can make a GET request on the job ID to retrieve the data for as long as the inference is on the server. The job ID is always returned when a document is sent successfully, it's important to store this ID when using webhooks for this type of scenario.

</details>

<details>

<summary><strong>How can I set up various environments like testing, staging, production?</strong></summary>

You can create any number of webhook endpoints: create one for each environment.

In your code, add an environment variable like `MINDEE_V2_WEBHOOK_ID` and set it according to the corresponding endpoint.

When sending a file for inference, [specify the webhook ID](/integrations/client-libraries-sdk/send-a-file-or-url#webhook-configuration) using the environment variable.

</details>

<details>

<summary><strong>I need to filter incoming requests, do you have a static IP?</strong></summary>

Our outgoing webhook server (incoming for you) does have a static IP address.

Enterprise customers can contact us for the IP range.

</details>


# Error Handling

Common error messages and their handling.

## Error Responses

All errors will contain the following information:

* Status: the HTTP status code.
* Title: a human-readable short description, it may not be unique.
* Detail: a human-readable description of the error, it will be unique.
* Code: a code to identify the specific error.\
  Provide this code when asking for support.

## 4xx Errors - Client Errors

<table><thead><tr><th width="130">HTTP Status</th><th>Possible Reasons</th></tr></thead><tbody><tr><td>400</td><td><ul><li>A required request parameter is missing.</li></ul></td></tr><tr><td>401</td><td><ul><li>Invalid API key. Make sure the key starts with <code>md_</code> and is <a href="/pages/xn2Lxhj8hX0u9RQWsaiS">active</a>.</li><li>Make sure you are calling the V2 API, your code should match <a href="/pages/qlViJTXZGr9z6S46KVSt">the samples</a>.</li><li>API keys are not JWTs: do not include <code>Bearer</code> in your <code>Authentication</code> header.</li></ul><p>Always prefer using the provided SDKs if possible.</p></td></tr><tr><td>402</td><td><ul><li>An optional feature is not in your plan.</li><li>Your subscription is either not active or exceeded limits.</li></ul></td></tr><tr><td>403</td><td><ul><li>You don't have access to the requested resource.</li></ul></td></tr><tr><td>404</td><td><ul><li>The requested resource does not exist.</li><li>The requested resource is not ready for use (usually a processing result).</li></ul></td></tr><tr><td>422</td><td><ul><li>Wrong format for a UUID.</li><li>Invalid or empty file sent.</li><li>Invalid parameter sent in a request.</li></ul></td></tr><tr><td>429</td><td><ul><li>Too many requests. Wait a few seconds and try again.<br>For more information, consult the section:  <a data-mention href="/pages/ikMUZj4uegWUNgcJJPD3#rate-limits">/pages/ikMUZj4uegWUNgcJJPD3#rate-limits</a>.</li></ul></td></tr></tbody></table>

## 5xx Errors - Server Errors

<table><thead><tr><th width="130">HTTP Status</th><th>Possible Reasons</th></tr></thead><tbody><tr><td>500</td><td><ul><li>Failed to run the inference.</li><li>Failed to process the request.</li></ul></td></tr></tbody></table>

### Mindee Status

Check to make sure that the API is fully operational.

You can also subscribe to be notified automatically.

<https://status.mindee.com/>


# Plans and Credits

Pricing plans, credits, and features comparison.

Mindee offers flexible subscription plans to suit individual developers, growing teams, and enterprise-scale organizations.

Choose the plan that best fits your support and automation needs.

Choose the number of credits that fit your volume.

**All plans** include access to:

* an unlimited number of all types of models ([Extraction](/extraction-models/extraction-models-overview), [Split](/split-models/split), [Crop](/crop-models/crop), etc)
* our AI agent to help you build your custom models and answer your questions
* our complete catalog of model templates to get started quickly
* the Mindee API for your integration

## Plan Overview & Pricing

<table><thead><tr><th width="156.1998291015625">Plan</th><th width="276.4000244140625">Monthly price (with annual billing)</th><th>Minimum Annual Credits</th></tr></thead><tbody><tr><td>Starter</td><td>€44 / month</td><td>6,000 credits</td></tr><tr><td>Pro</td><td>€116 / month</td><td>6,000 credits</td></tr><tr><td>Enterprise</td><td>Custom pricing</td><td>500,000+ credits</td></tr></tbody></table>

* All plans are available in **monthly** or **annual** billing.
* **Annual plans save 10%** compared to monthly pricing.
* You can **upgrade at any time**. Downgrades take effect at the end of your billing cycle.
* Overage charges are calculated automatically based on your plan’s rate.

More information is available in the section: [Billing](/account-management/billing).

### Credits

All plans have a minimum number of credits, however you may include more credits in your subscription to take advantage of lower credit prices.

Think of the plan as the features available, and the credits as the amount of documents to process.

## Try Mindee for Free!

You can try Mindee free for 14 days, this includes 200 pages. You can use the platform and make API calls.

Simply sign up to get started.

<a href="https://app.mindee.com/signup?utm_source=docs" class="button primary" data-icon="user-plus">Sign Up to Mindee</a>

Once the trial period ends, you’ll need to choose a subscription plan to continue using the service.

{% hint style="success" %}
During the trial period, we can activate certain features for you to test specific use cases.\
Don't hesitate to reach out to us!
{% endhint %}

## Select the Right Plan <a href="#h_bbb10a2ad9" id="h_bbb10a2ad9"></a>

When choosing between plans, it's essential to align your selection with your processing volume. Startups with minimal daily usage can start small, while enterprises with consistent heavy usage should consider higher-tier plans for scalability and technical support.

For example, the Starter Plan is ideal for businesses with limited API requirements, while teams handling higher daily volumes might benefit from the Business Plan.

Start with the Starter Plan if you're processing fewer than 100 pages per day. For higher API demands, such as 300 pages daily, the Business Plan can provide value and cost predictability.

## Feature Comparison

Some features are only applicable to [Extraction models](/extraction-models/extraction-models-overview).

| Feature                                                  |        Starter       |          Pro         |      Enterprise      |
| -------------------------------------------------------- | :------------------: | :------------------: | :------------------: |
| [AI Agent for custom model](#user-content-fn-1)[^1]      | :white\_check\_mark: | :white\_check\_mark: | :white\_check\_mark: |
| [Community support](#user-content-fn-2)[^2]              | :white\_check\_mark: | :white\_check\_mark: | :white\_check\_mark: |
| [Raw Text](#user-content-fn-3)[^3]                       | :white\_check\_mark: | :white\_check\_mark: | :white\_check\_mark: |
| RAG[^4]                                                  | :white\_check\_mark: | :white\_check\_mark: | :white\_check\_mark: |
| Polygons[^5]                                             | :white\_check\_mark: | :white\_check\_mark: | :white\_check\_mark: |
| [Confidence score](#user-content-fn-6)[^6]               | :white\_check\_mark: | :white\_check\_mark: | :white\_check\_mark: |
| [Boosted accuracy and precision](#user-content-fn-7)[^7] | :white\_check\_mark: | :white\_check\_mark: | :white\_check\_mark: |
| [Zero-retention controls](#user-content-fn-8)[^8]        | :white\_check\_mark: | :white\_check\_mark: | :white\_check\_mark: |
| Members[^9]                                              |          :x:         | :white\_check\_mark: | :white\_check\_mark: |
| [Data Processing Zone](#user-content-fn-10)[^10]         |          :x:         | :white\_check\_mark: | :white\_check\_mark: |
| [Support team](#user-content-fn-11)[^11]                 |          :x:         | :white\_check\_mark: | :white\_check\_mark: |
| Custom pricing                                           |          :x:         |          :x:         | :white\_check\_mark: |
| Dedicated account manager                                |          :x:         |          :x:         | :white\_check\_mark: |
| Premium technical support                                |          :x:         |          :x:         | :white\_check\_mark: |

## Plan Details

<table data-view="cards" data-full-width="false"><thead><tr><th></th><th></th><th data-hidden data-card-cover data-type="image">Cover image</th></tr></thead><tbody><tr><td><strong>Free Trial</strong></td><td><ul><li>It's free!</li><li>200 credits or 14 days</li><li>10 pages per document</li><li>Access to all model types</li><li>Access to <a href="/pages/FPe6NYrgvDdQ5xXFHVR7">Optional Features</a></li><li>API integration</li><li>AI Assistant for models</li><li>Docs Assistant </li></ul></td><td><a href="/files/HOim63cDgNIuonLNX94Q">/files/HOim63cDgNIuonLNX94Q</a></td></tr><tr><td><strong>Starter</strong></td><td><ul><li>Access to all model types</li><li>Unlimited models</li><li>Community support</li><li><a href="/pages/6oWHDmPeLdWNX58cCJPp">Raw Text</a> available</li><li><a href="/pages/uPDygoYr3mDz2whFnftm">Polygons</a> available</li><li><a href="/pages/zQlMhrvZxpl6scClGPRK">Confidence score</a> available</li><li><a href="/pages/zQlMhrvZxpl6scClGPRK">Boosted accuracy</a> and precision available</li><li>API integration</li><li>AI Assistant for models</li><li>Docs Assistant </li></ul></td><td><a href="/files/qy4a72TJILFTJ96UUm4d">/files/qy4a72TJILFTJ96UUm4d</a></td></tr><tr><td><mark style="color:yellow;"><strong>Pro</strong></mark></td><td><ul><li>Everything in Starter</li><li>Chat support</li><li>Members: multi-user management</li><li>data processing localization</li></ul></td><td><a href="/files/k8VYKbfRQjXmgCDwerDD">/files/k8VYKbfRQjXmgCDwerDD</a></td></tr><tr><td><mark style="color:red;"><strong>Enterprise</strong></mark></td><td><ul><li>Everything in Pro</li><li>Tailored usage and pricing</li><li>SLA-backed support</li><li>Dedicated account manager</li><li>Premium on-boarding and technical support</li></ul></td><td><a href="/files/sh7GqxdKT7GjQFLw2uCg">/files/sh7GqxdKT7GjQFLw2uCg</a></td></tr></tbody></table>

<a href="https://app.mindee.com/signup?utm_source=docs" class="button primary" data-icon="user-plus">Sign Up to Mindee</a>

## Frequently Asked Questions

<details>

<summary><strong>How is credit use calculated?</strong></summary>

Each file processed by Mindee will consume credits based on the number of pages contained in the uploaded file. Only files successfully processed count towards your credit consumption.

The exact amount of credits used per page is available for each model. For details, consult the section: [Credit Cost](/models/pricing).

Overall credit consumption can be tracked from a single dashboard. For details, consult the section: [Insights](/account-management/insights).

</details>

<details>

<summary><strong>Are there limits on the number of credits that can be consumed?</strong></summary>

Regardless of your purchased credit volume, there are no limits to credits you can consume.

This avoids any situations where you would be "locked" and unable to process more documents.

</details>

<details>

<summary><strong>Are there any differences between annual and monthly plans?</strong></summary>

Aside from the lower price, annual subscriptions provide more flexibility in credit usage.

For **annual plans**, all included credits are available at any time during the subscription year. Meaning all available credits can be used as you wish throughout the year. If you use more, they are billed as additional usage.\
​\
For **monthly plans**, credits are available each month. If you use less, they are lost; if you use more, they are billed as additional usage.

</details>

<details>

<summary><strong>Can I add features from another plan (à la carte pricing)?</strong></summary>

You cannot add features from another plan.

To benefit from a particular feature, you will need to upgrade to the plan which includes it.

</details>

<details>

<summary><strong>What happens to my data if I cancel my subscription?</strong></summary>

If you cancel your subscription, your data will be retained for 30 days to allow for potential reactivation or export. After this grace period, your documents and models will be automatically deleted, in accordance with our data lifecycle policy.

</details>

<details>

<summary><strong>Can the Mindee team add a feature I need?</strong></summary>

Our product team is always interested in adding new features to help our users.\
[Make a feature request!](https://feedback.mindee.com/?b=682f69c9e2404756e7e68d1c)

Enterprise customers can ask us directly for specific features they need.

</details>

[^1]: Build your own document parsers by describing the fields you need

[^2]: Access to our [Feedback](https://feedback.mindee.com/) page

[^3]: Extract the [full text content](/extraction-models/optional-features/raw-text-full-ocr) of your documents.

[^4]: Enhance your AI models with custom document knowledge bases.

    Learn more on our [RAG documentation](https://docs.mindee.com/models/improving-accuracy)

[^5]: Use advanced [polygon](/extraction-models/optional-features/polygons-bounding-boxes) extraction to capture data from complex layouts.

[^6]: Leverage [confidence scores](/extraction-models/optional-features/automation-confidence-score) to selectively automate document processing. Also boosts accuracy.

[^7]: Uses multiple models to [boost accuracy](/extraction-models/optional-features/automation-confidence-score). Also adds confidence scores.

[^8]: Storage policies to enable zero retention.

[^9]: Add members to your organization

[^10]: Choose the region where your data is processed: Europe, United States, or Global

[^11]: Get help from our support team via chat and priority support for Enterprise


# Support Levels

Mindee provides different levels of support depending on your subscription plan. Support options range from community-driven help to real-time chat with priority handling for urgent needs.

## Support Overview <a href="#support-overview" id="support-overview"></a>

Support level by [Plan](/account-management/plans).

<table><thead><tr><th width="264.5999755859375">Support</th><th width="151.199951171875" align="center">Starter</th><th align="center">Pro</th><th align="center">Business</th></tr></thead><tbody><tr><td><strong>Community Support</strong> <a href="https://feedback.mindee.com/">feedback.mindee.com</a></td><td align="center">✅</td><td align="center">✅</td><td align="center">✅</td></tr><tr><td><strong>Chat Support</strong></td><td align="center">❌</td><td align="center"><p>✅</p><p>Dedicated chat</p></td><td align="center"><p>✅</p><p>Dedicated chat</p></td></tr><tr><td><strong>Priority Support</strong></td><td align="center">❌</td><td align="center">❌</td><td align="center"><p>✅</p><p>Priority handling in chat</p></td></tr></tbody></table>

## Support Options <a href="#support-options" id="support-options"></a>

### Community Support

All plans include access to our community forum at [feedback.mindee.com](https://feedback.mindee.com/).\
This is the best place to ask questions, exchange ideas with other users, and access responses from the Mindee team.

### Chat Support

Available from the **Pro plan** onward, our Chat connects you directly with our support team via a dedicated chat channel for faster and more direct assistance.

### Priority Support

Included in the **Business plan**, Priority Support ensures that your requests are addressed first in the chat queue, reducing waiting times for critical issues.

## Choosing the Right Plan <a href="#choosing-the-right-level" id="choosing-the-right-level"></a>

* **Starter** ⇒ For individuals or early-stage projects where community-driven answers are sufficient.
* **Pro** ⇒ For teams that benefit from direct access to chat-based support.
* **Business** ⇒ For organizations where faster resolution and priority handling are required.


# Account Settings

Manage your Mindee account.

A Mindee account is required for logging into the platform.

However, models and API keys are stored at the organization level.

When you create your Mindee account, a default organization will be created. Your account will be the [owner](/account-management/organizations#member-roles) of this organization.

You can be a member of multiple organizations.

## Access Account Settings

Click on "<i class="fa-gear">:gear:</i> Settings" in the main menu, then on the "Account" tab.

Or simply click here: <a href="https://app.mindee.com/settings?tab=account" class="button primary">Go to Account page</a>

## Personal Information

### Email

If you connect to Mindee using your email and a password, you can change your email address. You'll be sent a link to confirm the new email address.

If you connect to Mindee through a 3rd party, such as Google or Github, the email address used will be displayed. It is not possible to change it, you'll need to connect with another provider or use an email/password login.

### Full Name

You can change your full name as it appears in the Mindee platform to other users in your organization.

## Danger Zone

### Delete Account

You can delete your account and all its associated data.

{% hint style="danger" %}
Deleting your account is an **irreversible action**, all data will be lost forever!
{% endhint %}

Deleting your account will also delete all organizations in which you are the only owner.

Any non-owner users will lose access to these organizations.


# Organization & Members

Organizations in Mindee enable collaborative work: users can belong to an organization and participate in workspaces with role-based access control.

## Access Organization Settings

Click on "<i class="fa-gear">:gear:</i> Settings" in the main menu, then on the "Organization" tab.

Or simply click here: <a href="https://app.mindee.com/settings?tab=organization" class="button primary">Go to Organization page</a>

## Organization Details

Customize the display name of your organization.

Access your Organization ID, a unique identifier generated automatically.

### Default Processing Zone

Set the processing zone used for all models by default.

Most users will want to set the processing zone here rather than at the model level.

## Actions

### Usage Alerts

You can set an alert to receive an email when your [plan](/account-management/plans) has reached a certain percentage of use.

## **Team Members**

Each organization may have several members.

Each user ([account](/account-management/plans)) may be a member of several organizations.

The **Team Members** section lists all users who belong to your organization, along with their:

* **Name & Email** – displayed for clarity.
* **Role** – "Owner", "Administrator", or "Member"
* **Status** – shows as *Active* for members with access.

### Add New Team Members

Use the "Invite Member" button to add a new user.

Enter their email address and select their role.

They will receive an email allowing them to create an account and join your organization.

{% hint style="info" icon="money-check-dollar-pen" %}
You can only invite members to your organization if you have the the Pro, Business, or Enterprise [Plans and Credits](/account-management/plans).
{% endhint %}

### Remove Team Members

Use the "<i class="fa-trash-can">:trash-can:</i>" button to remove a member from your organization.

You must have the "Administrator" or "Owner" role to remove users.

You must have the "Owner" role to remove users with the "Administrator" role.

### Update a Member's Role

Use the "Edit Role" button to change a member's role.

You must have the "Administrator" or "Owner" role to update users.

You must have the "Owner" role to update users with the"Administrator" role.

#### Member Roles

Each organization is governed by a **role system**.

A user must be explicitly added to an organization to access it, except for the **creator**, who is automatically granted the "Owner" role.

<table><thead><tr><th width="168.9998779296875">Role Name</th><th>Permissions Summary</th></tr></thead><tbody><tr><td>Member</td><td>Add and manage models<br>Run Live Tests<br>Manage API keys</td></tr><tr><td>Administrator</td><td><p>All "Member" permissions<br>Manage "Member" users</p><p>Update organization settings<br>Manage billing setting</p></td></tr><tr><td>Owner</td><td><p>All "Administrator" permissions</p><p>Can add and remove "Administrator" users</p></td></tr></tbody></table>

{% hint style="info" %}
Only one "Owner" per organization. This owner is set at creation and cannot be transferred.
{% endhint %}

## Danger Zone

### Delete Organization

You can delete your organization and all its associated data.

{% hint style="danger" %}
Deleting your organization is an **irreversible action**, all data will be lost forever!
{% endhint %}

You must always belong to at least one organization.

Deleting an organization will not delete user accounts, this includes your account and any other users belonging to the organization.

When you have a single organization in which you are the sole owner:

* To start with a "fresh" organization: first create the new organization, then delete the old one.
* To delete all your Mindee data: [delete your account](/account-management/account-settings#delete-account) instead, the organization will be automatically deleted.


# Billing

The Billing page in your Mindee account lets you manage your subscription and access past invoices.

## Overview

From the Billing page, you can:

* View your current plan
* Choose a new plan (upgrade or downgrade)
* Access your Stripe customer portal

## Access the Billing Page

1. Go to [app.mindee.com](https://app.mindee.com/)
2. On the left-hand menu, click on "<i class="fa-gear">:gear:</i> Settings"
3. Select the "Billing" tab

Or simply click here: <a href="https://app.mindee.com/settings?tab=billing" class="button primary">Go to Billing page</a>

Only [users](/account-management/organizations#team-members) of type "Administrator" or "Owner" may access and modify billing details.

## Change Your Plan

To change your plan:

1. Find the plan under the available options
2. Click on the "Downgrade" or "Upgrade" button at the bottom of the plan

You can upgrade to a higher plan at any time. Your current plan's commercial terms, including the remaining subscription period and pricing, will be taken into account when switching to the new plan.

## Stripe Customer Portal

To access your Stripe customer portal, click on the "Manage Subscription" button.

<figure><img src="/files/89kSt8gxqJvmTlNKYbDn" alt="Manage Subscription button - access Stripe customer portal" width="563"><figcaption></figcaption></figure>

### Access Invoices

{% hint style="info" %}
Invoices are issued automatically at the start of each billing period.
{% endhint %}

Your past invoices are available in the "Invoice History" section of your organization's billing page.

From there you can also pay any outstanding invoices and download payment receipts.

**Note**: only "Administrator" and "Owner" members may access the Invoice History section, for other members it is not present on the page at all.

Additionally, you can download past invoices or pay any outstanding invoices from your Stripe account.

### Set Your Tax ID

You can set your Tax ID at any time. After clicking on "Manage Subscription", click on "Update Information", and finally edit the "Tax ID" field.

## Supported Payment Methods

Any payment method supported by Stripe can be used.

Among these, the most popular are credit and debit cards.

Custom payment options are available for Enterprise customers.

## Included Credits and Overages

Each plan includes monthly or yearly included credits depending on your subscription frequency. There is no service disruption if you exceed the included credits.

If you exceed the included credits in your plan, these additional credits will be charged according to the Overage Rate. These extra credits will be invoiced every €100, or at the end of your subscription.

| Plan    | Included Credits | Overage Rate                 |
| ------- | ---------------- | ---------------------------- |
| Starter | 6,000 per year   | €0.044 per additional credit |
| Pro     | 6,000 per year   | €0.044 per additional credit |

{% hint style="info" %}
**Pages are calculated based on successfully processed documents.**

Client errors (HTTP 4xx) and processing errors will **not** count towards your page usage.
{% endhint %}

For details on tracking your page and credit usage, consult the section: [Insights](/account-management/insights).

Additionally, API responses contain the number of pages processed. Look in the section: [Response Processing](/integrations/client-libraries-sdk/process-the-response#file-metadata).

Finally, you can be notified when your plan is reaching its limit. For details, check the section: [Organization & Members](/account-management/organizations#usage-alerts).

## Billing Frequency

You will be billed as soon as you complete your subscription payment online.

Payment is required upfront to access the platform and the features included in your plan.

Enterprise customers have access to different billing methods, please reach out to us directly.

## Frequently Asked Questions

<details>

<summary><strong>How do additional credit charges appear on the invoice?</strong></summary>

Additional credit usage is itemized separately, on the invoice as "overage charges".

The invoice will show:

* the base cost of your selected [plan](/account-management/plans)
* any additional credits consumed beyond your plan's included credits
* cost per credit

</details>


# Insights

The insights page allows you to view various data concerning your usage of the Mindee platform and API.

## Overview

Use the insights page to view all of your requests over a given period.

Notably, you can keep track of your credit usage.

## Access the Insights Page

1. Go to [app.mindee.com](https://app.mindee.com/)
2. On the left-hand menu, click on **Insights**

Or simply click here: <a href="https://app.mindee.com/insights" class="button primary">Go to Insights page</a>

## Data Views

Selecting the Data View is done at the top of the page:

<div align="left"><figure><img src="/files/ZQJXuydj9LQyDYhfyr7Z" alt="Insights - data view selector"><figcaption></figcaption></figure></div>

All Data Views will show the total usage, at the bottom of the page:

<figure><img src="/files/84l03E0JEgi3gQGQs5OS" alt=""><figcaption></figcaption></figure>

### Traffic View

The Traffic view shows the number of calls and the number of pages that were successfully processed.

These requests count towards your plan's credit usage, the number of credits consumed is shown.

This does not include calls that resulted in an error.

<figure><img src="/files/6UtPlswZmMOSZbfATSug" alt="insights - traffic shows requests, pages, credits" width="563"><figcaption></figcaption></figure>

### Processing Time View

The Processing Time view shows the average time requests take to process, in seconds.

Processing time is calculated from when the request was received to when the [inference](/getting-started/glossary) was finished.

<figure><img src="/files/jAw0fBjVDfMY6WRLkif1" alt="insights - processing time" width="563"><figcaption></figcaption></figure>

### Errors View

The Errors view shows the number of processing ([inference](/getting-started/glossary#inference)) errors.

This does not include user errors (HTTP 4xx), since these requests get rejected before any processing takes place.

Errors are shown for informational purposes only, they do not count towards your plan's credit usage.

<figure><img src="/files/i23AY5b5xNk8uuVqk7z4" alt="insights - errors" width="563"><figcaption></figcaption></figure>

## Data Filters

For all data views, the following filters are available.

### Origin Filter

Filter based on the origin of the request.

<div align="left"><figure><img src="/files/cJaJW9tde5kC9Da4niMg" alt="Insights - filter data based on origin"><figcaption></figcaption></figure></div>

"Live Test" is when using the [live test](/models/live-test) functionality on the platform.

"API" is for any request made using an API key.

### Date Range Filter

Filter based on start and end dates.

<div align="left"><figure><img src="/files/isCcdifdMOPqf66vPjit" alt=""><figcaption></figcaption></figure></div>

### Group By

Only affects the visualization. Group results by day, month, or year.

<div align="left"><figure><img src="/files/pNKmsvu6OfWJbc46lYeD" alt=""><figcaption></figcaption></figure></div>

### API Key Filter

Show all API keys or filter on a specific one.

<div align="left"><figure><img src="/files/YJBvBaLsIrjrWYzQ4pnP" alt=""><figcaption></figcaption></figure></div>

### Activated Options

Filter by [optional features](/extraction-models/optional-features) activated in the call.

<div align="left"><figure><img src="/files/zeeifn8EVtlwZCIag2Jf" alt=""><figcaption></figcaption></figure></div>

3 states to choose from:

* empty: no filter
* check mark: show only calls with the option **active**
* minus sign: show only calls with the option **inactive**


# Extraction Use Cases

Common document extraction models you can build with Mindee.

Use Mindee to extract structured data from many document types.\
You can start from a template or build a custom model.

* **Start from the Catalog:** pick a ready-made model template (Invoice, Passport, etc.) and adapt its Data Schema to your needs.
* **Build a custom model:** create a custom model from scratch for documents unique to your needs.

Here are some examples of models you can build with Mindee:

## Finance & Business

* [**Invoice**](/use-cases/extraction-models/invoice)**:** Extract key data like supplier, amounts, due dates, and line items.
* [**Receipt**](/use-cases/extraction-models/receipt)**:** Capture totals, taxes, and line items for expense automation.
* [**Financial Document**](/use-cases/extraction-models/financial-documents)**:** Handle mixed inputs such as invoices or receipts in one workflow.
* [**Bank Statement**](/use-cases/extraction-models/bank-statement)**:** Extract account numbers, statement dates, and transaction data from bank statements.
* [**Payslip**](/use-cases/extraction-models/payslip)**:** Extract salary, employer, pay period, and deduction data from payslips.
* [**Bank Account Details**](/use-cases/extraction-models/bank-account-details)**:** Capture account-holder, IBAN, and banking details for onboarding or KYC flows.

## Identity & Verification

* [**International ID Card**](/use-cases/extraction-models/international-id-card)**:** Capture personal details for KYC or onboarding.
* [**Passport**](/use-cases/extraction-models/passport)**:** Parse official ID data from passports with high accuracy.
* [**Driver's License**](/use-cases/extraction-models/drivers-license)**:** Extract identity details from driver licenses for fast verification.
* [**Business Card**](/use-cases/extraction-models/business-card)**:** Extract names, job titles, email addresses, phone numbers, and company details.
* [**US Healthcare Card**](/use-cases/extraction-models/us-healthcare-card)**:** Capture member, provider, and plan details from US healthcare cards.
* [**European Vehicle Registration**](/use-cases/extraction-models/european-vehicle-registration)**:** Extract registration numbers, VINs, and owner details from EU vehicle registration documents.

## Travel & Logistics

* [**Boarding Pass**](/use-cases/extraction-models/boarding-pass)**:** Read passenger and flight details directly from airline passes.
* [**Bill of Lading**](/use-cases/extraction-models/bill-of-lading)**:** Extract key data like shipper, consignee, and items shipped.

## Labels & Packaging

* [**Nutrition Facts**](/use-cases/extraction-models/nutrition-facts)**:** Extract serving sizes, calories, and nutrient values from product labels.


# Invoice

Automatically parse invoices and extract structured financial data using the Invoice template available in the model Catalog.

Watch our quick demo to see how you can easily create your custom invoice model with Mindee:

{% @supademo/embed url="<https://app.supademo.com/demo/cmfci0aiy7nt739ozcvdq9hqa>" demoId="cmfci0aiy7nt739ozcvdq9hqa" %}

## Why use Mindee for Invoices?

Invoices can vary widely by supplier, format, and layout. The Invoice model template is designed to handle these differences, so you get consistent and reliable data extraction without custom development.

Common use cases:

* Accounts payable automation
* Expense and cost tracking
* Tax and compliance reporting

## Two Ways to Start Building your Invoice Model

### 1. Choose "Invoice" in the Catalog (Recommended)

* Click on "Create your document AI model" in your dashboard, then select **"Invoice".**
* The Invoice model template comes pre-configured with the standard [Invoice](/use-cases/extraction-models/invoice#invoice-fields).
* Once your Invoice model is created, you can immediately [test](/models/live-test) with your own invoices.
* Optionally, you can adjust the model's [Data Schema](/extraction-models/data-schema) if you need to modify fields.

### 2. Build a Custom Invoice Model with the AI Agent

* If your workflow requires extra fields (e.g. purchase order number, IBAN, payment terms), you can describe them directly to the AI Agent.
* Optionally upload a sample invoice for context.
* The Agent will generate a tailored schema that extends beyond the built-in Invoice parser.

You can use this invoice sample if you want to try and do a live test yourself:

<figure><img src="/files/R1HwyW4ID455JL3rkQTR" alt="a fake invoice from Turnpike Designs" width="533"><figcaption></figcaption></figure>

## Document format support

The Invoice model accepts PDFs and common image formats (JPG, PNG). It works reliably with scanned, photographed, and digital invoices. Like all Mindee catalog models, handwriting can be recognized in addition to printed text.

## Invoice Fields

Documentation for all fields present in the data schema.

Field *accessors* are used as the keys for accessing the values in the returned data. On the data schema interface, this is the "Field Name".

Field *value types* indicate how the value is returned by the API. On the data schema interface, this is the "Field Type".

<details>

<summary>Supplier Name</summary>

The name of the supplier of the invoice.

Accessor: `supplier_name`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Supplier Phone Number</summary>

The phone number of the supplier of the invoice.

Accessor: `supplier_phone_number`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Customer Company Registration</summary>

A list of company registration details including type and number, for the customer of the invoice.

Accessor: `customer_company_registration`

**Subfields**

* **Number**\
  The company registration number.\
  Accessor: `number`\
  Value Type: `string`
* **Type**\
  The type of the company registration number.\
  Accessor: `type`\
  Possible Values: `VAT`, `SIRET`, `SIREN`, `NIF`, `CF`, `UID`, `STNR`, `HRA_HRB`, `TIN`, `RFC`, `BTW`, `ABN`, `UEN`, `CVR`, `ORGNRO`, `INN`, `DPH`, `NIP`, `GSTIN`, `CRN`, `KVK`, `DIC`, `TAX_ID`, `CIF`, `GST_HST_CA`, `COC`

Can have multiple values (is a list/array).

</details>

<details>

<summary>Supplier Company Registration</summary>

A list of company registration details including type and number, for the supplier of the invoice.

Accessor: `supplier_company_registration`

**Subfields**

* **Number**\
  The company registration number.\
  Accessor: `number`\
  Value Type: `string`
* **Type**\
  The type of the company registration number.\
  Accessor: `type`\
  Possible Values: `VAT`, `SIRET`, `SIREN`, `NIF`, `CF`, `UID`, `STNR`, `HRA_HRB`, `TIN`, `RFC`, `BTW`, `ABN`, `UEN`, `CVR`, `ORGNRO`, `INN`, `DPH`, `NIP`, `GSTIN`, `CRN`, `KVK`, `DIC`, `TAX_ID`, `CIF`, `GST_HST_CA`, `COC`

Can have multiple values (is a list/array).

</details>

<details>

<summary>Invoice Number</summary>

The number of the invoice.

Accessor: `invoice_number`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Date</summary>

The date the invoice was issued.

Accessor: `date`\
Value Type: `date`

Has a single value.

</details>

<details>

<summary>Total Amount</summary>

The final total amount paid, including all taxes and discounts.

Accessor: `total_amount`\
Value Type: `number`

Has a single value.

</details>

<details>

<summary>Total Net</summary>

The total amount before taxes.

Accessor: `total_net`\
Value Type: `number`

Has a single value.

</details>

<details>

<summary>Total Tax</summary>

The total amount of all taxes.

Accessor: `total_tax`\
Value Type: `number`

Has a single value.

</details>

<details>

<summary>Taxes</summary>

A list of individual taxes applied, each including rate, base and amount.

Accessor: `taxes`

**Subfields**

* **Rate**\
  The tax rate of this tax as a decimal.\
  Accessor: `rate`\
  Value Type: `number`
* **Base**\
  The base amount on which this tax is computed.\
  Accessor: `base`\
  Value Type: `number`
* **Amount**\
  The computed tax amount for this tax.\
  Accessor: `amount`\
  Value Type: `number`

Can have multiple values (is a list/array).

</details>

<details>

<summary>Line Items</summary>

A list of line items in the invoice.

Accessor: `line_items`

**Subfields**

* **Description**\
  A description of the item or service.\
  Accessor: `description`\
  Value Type: `string`
* **Quantity**\
  The quantity of the item or service.\
  Accessor: `quantity`\
  Value Type: `number`
* **Unit Price**\
  The price per unit of the item or service.\
  Accessor: `unit_price`\
  Value Type: `number`
* **Total Price**\
  The total price for the line item: quantity \* unit price.\
  Accessor: `total_price`\
  Value Type: `number`
* **Product Code**\
  The product code of the item.\
  Accessor: `product_code`\
  Value Type: `string`
* **Tax Amount**\
  The tax amount of the item.\
  Accessor: `tax_amount`\
  Value Type: `number`
* **Tax Rate**\
  The tax rate of the item.\
  Accessor: `tax_rate`\
  Value Type: `number`
* **Unit Measure**\
  The unit of measure of the item.\
  Accessor: `unit_measure`\
  Value Type: `string`

Can have multiple values (is a list/array).

</details>

<details>

<summary>Document Type</summary>

Document type of the financial document.

Accessor: `document_type`\
Possible Values: `invoice`, `payslip`, `quote`, `purchase_order`, `statement`, `receipt`, `credit_note`, `other_financial`

Has a single value.

</details>

<details>

<summary>Locale</summary>

The locale contains the language, country and currency of the invoice.

Accessor: `locale`

**Subfields**

* **Language**\
  The language of the invoice, ISO 639-1 language code.\
  Accessor: `language`\
  Value Type: `string`
* **Country**\
  The country of the invoice, ISO 3166-1 alpha-2.\
  Accessor: `country`\
  Value Type: `string`
* **Currency**\
  The currency in which is issued the invoice, ISO 4217.\
  Accessor: `currency`\
  Value Type: `string`

Has a single value.

</details>

<details>

<summary>Customer Name</summary>

The name of the customer of the invoice.

Accessor: `customer_name`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Customer Address</summary>

The full address of the customer and the breakdown of the address in components.

Accessor: `customer_address`

**Subfields**

* **Address**\
  The full raw address of the customer as written in the document.\
  Accessor: `address`\
  Value Type: `string`
* **Street number**\
  The street number of the address.\
  Accessor: `street_number`\
  Value Type: `string`
* **Street Name**\
  The street name of the address.\
  Accessor: `street_name`\
  Value Type: `string`
* **PO Box**\
  PO Box number, if there is one in the address.\
  Accessor: `po_box`\
  Value Type: `string`
* **Address Complement**\
  Address complement: floor, building, suite, ...\
  Accessor: `address_complement`\
  Value Type: `string`
* **City**\
  The city of the address.\
  Accessor: `city`\
  Value Type: `string`
* **Postal Code**\
  The postal code of the address.\
  Accessor: `postal_code`\
  Value Type: `string`
* **State**\
  The state / region / land of the address, if there is any.\
  Accessor: `state`\
  Value Type: `string`
* **Country**\
  The country of the address.\
  Accessor: `country`\
  Value Type: `string`

Has a single value.

</details>

<details>

<summary>Shipping Address</summary>

The full shipping address and the breakdown of the address in components.

Accessor: `shipping_address`

**Subfields**

* **Address**\
  The full raw shipping address as written in the document.\
  Accessor: `address`\
  Value Type: `string`
* **Street number**\
  The street number of the address.\
  Accessor: `street_number`\
  Value Type: `string`
* **Street Name**\
  The street name of the address.\
  Accessor: `street_name`\
  Value Type: `string`
* **PO Box**\
  PO Box number, if there is one in the address.\
  Accessor: `po_box`\
  Value Type: `string`
* **Address Complement**\
  Address complement: floor, building, suite, ...\
  Accessor: `address_complement`\
  Value Type: `string`
* **City**\
  The city of the address.\
  Accessor: `city`\
  Value Type: `string`
* **Postal Code**\
  The postal code of the address.\
  Accessor: `postal_code`\
  Value Type: `string`
* **State**\
  The state / region / land of the address, if there is any.\
  Accessor: `state`\
  Value Type: `string`
* **Country**\
  The country of the address.\
  Accessor: `country`\
  Value Type: `string`

Has a single value.

</details>

<details>

<summary>Billing Address</summary>

The full billing address and the breakdown of the address in components.

Accessor: `billing_address`

**Subfields**

* **Address**\
  The full full raw billing address as written in the document.\
  Accessor: `address`\
  Value Type: `string`
* **Street number**\
  The street number of the address.\
  Accessor: `street_number`\
  Value Type: `string`
* **Street Name**\
  The street name of the address.\
  Accessor: `street_name`\
  Value Type: `string`
* **PO Box**\
  PO Box number, if there is one in the address.\
  Accessor: `po_box`\
  Value Type: `string`
* **Address Complement**\
  Address complement: floor, building, suite, ...\
  Accessor: `address_complement`\
  Value Type: `string`
* **City**\
  The city of the address.\
  Accessor: `city`\
  Value Type: `string`
* **Postal Code**\
  The postal code of the address.\
  Accessor: `postal_code`\
  Value Type: `string`
* **State**\
  The state / region / land of the address, if there is any.\
  Accessor: `state`\
  Value Type: `string`
* **Country**\
  The country of the address.\
  Accessor: `country`\
  Value Type: `string`

Has a single value.

</details>

<details>

<summary>Supplier Address</summary>

The full address of the supplier and the breakdown of the address in components.

Accessor: `supplier_address`

**Subfields**

* **Address**\
  The full raw supplier address as written in the document.\
  Accessor: `address`\
  Value Type: `string`
* **Street number**\
  The street number of the address.\
  Accessor: `street_number`\
  Value Type: `string`
* **Street Name**\
  The street name of the address.\
  Accessor: `street_name`\
  Value Type: `string`
* **PO Box**\
  PO Box number, if there is one in the address.\
  Accessor: `po_box`\
  Value Type: `string`
* **Address Complement**\
  Address complement: floor, building, suite, ...\
  Accessor: `address_complement`\
  Value Type: `string`
* **City**\
  The city of the address.\
  Accessor: `city`\
  Value Type: `string`
* **Postal Code**\
  The postal code of the address.\
  Accessor: `postal_code`\
  Value Type: `string`
* **State**\
  The state / region / land of the address, if there is any.\
  Accessor: `state`\
  Value Type: `string`
* **Country**\
  The country of the address.\
  Accessor: `country`\
  Value Type: `string`

Has a single value.

</details>

<details>

<summary>Due Date</summary>

The date on which the invoice is due.

Accessor: `due_date`\
Value Type: `date`

Has a single value.

</details>

<details>

<summary>PO Number</summary>

The purchase order number.

Accessor: `po_number`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Reference Numbers</summary>

List of all reference numbers on the invoice, including the purchase order number.

Accessor: `reference_numbers`\
Value Type: `string`

Can have multiple values (is a list/array).

</details>

<details>

<summary>Payment Date</summary>

The date on which the payment is due / was full-filled.

Accessor: `payment_date`\
Value Type: `date`

Has a single value.

</details>

<details>

<summary>Supplier Payment Details</summary>

List of payment details associated to the supplier of the invoice.

Accessor: `supplier_payment_details`

**Subfields**

* **IBAN**\
  The International Bank Account Number (IBAN).\
  Accessor: `iban`\
  Value Type: `string`
* **SWIFT**\
  The bank s SWIFT Business Identifier Code (BIC).\
  Accessor: `swift`\
  Value Type: `string`
* **Account Number**\
  The account number.\
  Accessor: `account_number`\
  Value Type: `string`
* **Routing Number**\
  The routing number.\
  Accessor: `routing_number`\
  Value Type: `string`

Can have multiple values (is a list/array).

</details>

<details>

<summary>Supplier Website</summary>

The website URL of the supplier or merchant.

Accessor: `supplier_website`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Supplier Email</summary>

The email address of the supplier or merchant.

Accessor: `supplier_email`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Customer ID</summary>

The customer account number or identifier from the supplier.

Accessor: `customer_id`\
Value Type: `string`

Has a single value.

</details>


# Receipt

Use our pre-trained Receipt model or adjust with the fields you need with Mindee V2.

Here is a quick demo of Mindee V2's Receipt model:

{% @supademo/embed url="<https://app.supademo.com/demo/cmiytadiq07vk14g4techyr3b>" demoId="cmiytadiq07vk14g4techyr3b" %}

## Why Use Mindee for Receipts?

Receipts vary immensely in format, country, language, and quality. Mindee simplifies extraction and ensures high reliability by enabling you to:

* Handle global formats: Our model is trained on receipts from over 50 countries, automatically extracting data points regardless of local layout or language.
* Process poor quality inputs: Robustly extracts data from scanned documents, mobile photos, and even handwritten text on certain fields.
* Capture detailed line items: Accurately extract complex, nested data like individual line items, quantities, and prices for granular expense tracking.
* Get structured output with zero configuration: Start instantly with a pre-trained model that extracts standard fields like total amount, date, vendor name, and expense category.

## Two Ways to Start Using the Receipt Model

Mindee offers two paths to start extracting data from receipts:

### 1. Choose "Receipt" in the Catalog (Recommended)

* Click on "Create your document AI model" in your dashboard, then select **"Receipt".**
* The Receipt model template comes pre-configured with standard [Receipt](/use-cases/extraction-models/receipt#receipt-fields).
* Once your Invoice model is created, you can immediately [test](/models/live-test) with your own invoices.
* Optionally, you can adjust the model's [Data Schema](/extraction-models/data-schema) if you need to modify fields.

### **2. Customize the Model via Data Schema**

You can instantly tailor the pre-trained model to your exact needs. By navigating to the Data Schema interface, you can:

* Add new fields that are unique to your documents (e.g., internal identifiers, specific customer IDs).
* Delete existing fields that are not relevant to your use case.
* Refine the extraction by modifying field types or adding custom instructions for the AI assistant.

## Receipt Fields

Documentation for all fields present in the data schema.

Field *accessors* are used as the keys for accessing the values in the returned data. On the data schema interface, this is the "Field Name".

Field *value types* indicate how the value is returned by the API. On the data schema interface, this is the "Field Type".

<details>

<summary>Supplier Name</summary>

The name of the supplier of the receipt.

Accessor: `supplier_name`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Supplier Address</summary>

The full address of the supplier who delivered the receipt.

Accessor: `supplier_address`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Supplier Phone Number</summary>

The phone number of the supplier of the receipt.

Accessor: `supplier_phone_number`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Supplier Company Registration</summary>

A list of company registration details including type and number, for the supplier of the receipt.

Accessor: `supplier_company_registration`

**Subfields**

* **Number**\
  The company registration number.\
  Accessor: `number`\
  Value Type: `string`
* **Type**\
  The type of the company registration number.\
  Accessor: `type`\
  Possible Values: `VAT`, `SIRET`, `SIREN`, `NIF`, `CF`, `UID`, `STNR`, `HRA_HRB`, `TIN`, `RFC`, `BTW`, `ABN`, `UEN`, `CVR`, `ORGNRO`, `INN`, `DPH`, `NIP`, `GSTIN`, `CRN`, `KVK`, `DIC`, `TAX_ID`, `CIF`, `GST_HST_CA`, `COC`

Can have multiple values (is a list/array).

</details>

<details>

<summary>Receipt Number</summary>

The number of the receipt.

Accessor: `receipt_number`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Date</summary>

The date the receipt was issued.

Accessor: `date`\
Value Type: `date`

Has a single value.

</details>

<details>

<summary>Time</summary>

The time the receipt was issued.

Accessor: `time`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Total Amount</summary>

The final total amount paid, including all taxes and discounts.

Accessor: `total_amount`\
Value Type: `number`

Has a single value.

</details>

<details>

<summary>Total Net</summary>

The total amount before taxes.

Accessor: `total_net`\
Value Type: `number`

Has a single value.

</details>

<details>

<summary>Total Tax</summary>

The total amount of all taxes.

Accessor: `total_tax`\
Value Type: `number`

Has a single value.

</details>

<details>

<summary>Taxes</summary>

A list of individual taxes applied, each including rate, base and amount.

Accessor: `taxes`

**Subfields**

* **Rate**\
  The tax rate of this tax as a decimal.\
  Accessor: `rate`\
  Value Type: `number`
* **Base**\
  The base amount on which this tax is computed.\
  Accessor: `base`\
  Value Type: `number`
* **Amount**\
  The computed tax amount for this tax.\
  Accessor: `amount`\
  Value Type: `number`

Can have multiple values (is a list/array).

</details>

<details>

<summary>Tips &#x26; Gratuity</summary>

The total tips and gratuities for the receipt.

Accessor: `tips_gratuity`\
Value Type: `number`

Has a single value.

</details>

<details>

<summary>Line Items</summary>

A list of line items in the receipt, each including description, quantity, unit price and total price.

Accessor: `line_items`

**Subfields**

* **Description**\
  A description of the item or service.\
  Accessor: `description`\
  Value Type: `string`
* **Quantity**\
  The quantity of the item or service.\
  Accessor: `quantity`\
  Value Type: `number`
* **Unit Price**\
  The price per unit of the item or service.\
  Accessor: `unit_price`\
  Value Type: `number`
* **Total Price**\
  The total price for the line item: quantity \* unit price.\
  Accessor: `total_price`\
  Value Type: `number`

Can have multiple values (is a list/array).

</details>

<details>

<summary>Document Type</summary>

Document type: expense receipt or credit card receipt

Accessor: `document_type`\
Possible Values: `expense_receipt`, `credit_card_receipt`

Has a single value.

</details>

<details>

<summary>Purchase Category</summary>

The category of the receipt.

Accessor: `purchase_category`\
Possible Values: `food`, `gasoline`, `parking`, `toll`, `accommodation`, `transport`, `telecom`, `software`, `shopping`, `energy`, `miscellaneous`

Has a single value.

</details>

<details>

<summary>Purchase Subcategory</summary>

The purchase subcategory of the receipt.

Accessor: `purchase_subcategory`\
Possible Values: `restaurant`, `delivery`, `train`, `public`, `taxi`, `car_rental`, `plane`, `micromobility`, `office_supplies`, `electronics`, `cultural`, `groceries`, `other`

Has a single value.

</details>

<details>

<summary>Locale</summary>

The locale contains the language, country and currency of the receipt.

Accessor: `locale`

**Subfields**

* **Language**\
  The language of the receipt, ISO 639-1 language code.\
  Accessor: `language`\
  Value Type: `string`
* **Country**\
  The country of the receipt, ISO 3166-1 alpha-2.\
  Accessor: `country`\
  Value Type: `string`
* **Currency**\
  The currency in which is issued the receipt, ISO 4217.\
  Accessor: `currency`\
  Value Type: `string`

Has a single value.

</details>


# Financial Document

Process a wide variety of financial documents and adapted to your specific needs.

Here is a quick demo on how to set up your Financial Document model:

{% @supademo/embed url="<https://app.supademo.com/demo/cmfdw3e3q90i539ozo0p7lqiu>" demoId="cmfdw3e3q90i539ozo0p7lqiu" %}

## Why Use Mindee for Financial Documents?

### Unified processing for multiple financial document types

With the Financial Document model, you don’t need to build a separate parser for each kind of document.\
Invoices, bank statements, receipts, or even balance sheets can all be uploaded to the same endpoint.\
The model automatically interprets the type of document and extracts the relevant fields accordingly.

This allows you to:

* Centralize all your financial document workflows in one place
* Avoid managing multiple APIs for each document type
* Scale faster when new document formats appear in your processes

### Common Use Cases

You can use Mindee to extract structured data from:

* Invoices (supplier/customer name, totals, tax lines, due dates, etc.)
* Bank statements (account holder, IBAN, balances)
* Payment confirmations and receipts
* Account summaries or balance sheets
* Custom internal financial reports

## Two Ways to Get Started

### 1. Choose "Financial Document" in the Catalog (Recommended)

* Click on "Create your document AI model" in your dashboard, then select **"Financial Document".**
* The Financial Document model template comes pre-configured with standard [#financial-document-fields](#financial-document-fields "mention").
* Once your Invoice model is created, you can immediately [test](/models/live-test) with your own invoices.
* Optionally, you can adjust the model's [Data Schema](/extraction-models/data-schema) if you need to modify fields.

### **2. Build a tailored Financial Document model with the AI Agent**

* If you need additional or non-standard fields (e.g. purchase order number, internal reference codes, bank account details), start a conversation with the Agent.
* Describe what you want extracted and optionally upload a sample document.
* The Agent will propose a schema, which you can refine until it matches your requirements.

You can use this financial document sample to do a live test yourself:

<figure><img src="/files/0xk4Sh1t9IKhnNcJInR1" alt="a fake invoice from John Smith" width="563"><figcaption></figcaption></figure>

## Supported Formats

* **PDF files:** single-page or multi-page
* **Images:** JPG, PNG, TIFF, and more

See full [list of accepted files](https://docs.mindee.com/integrations/technical-limitations#accepted-files).

## Financial Document Fields

Documentation for all fields present in the data schema.

Field *accessors* are used as the keys for accessing the values in the returned data. On the data schema interface, this is the "Field Name".

Field *value types* indicate how the value is returned by the API. On the data schema interface, this is the "Field Type".

<details>

<summary>Supplier Name</summary>

The name of the supplier of the invoice.

Accessor: `supplier_name`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Supplier Phone Number</summary>

The phone number of the supplier of the invoice.

Accessor: `supplier_phone_number`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Customer Company Registration</summary>

A list of company registration details including type and number, for the customer of the invoice.

Accessor: `customer_company_registration`

**Subfields**

* **Number**\
  The company registration number.\
  Accessor: `number`\
  Value Type: `string`
* **Type**\
  The type of the company registration number.\
  Accessor: `type`\
  Possible Values: `VAT`, `SIRET`, `SIREN`, `NIF`, `CF`, `UID`, `STNR`, `HRA_HRB`, `TIN`, `RFC`, `BTW`, `ABN`, `UEN`, `CVR`, `ORGNRO`, `INN`, `DPH`, `NIP`, `GSTIN`, `CRN`, `KVK`, `DIC`, `TAX_ID`, `CIF`, `GST_HST_CA`, `COC`

Can have multiple values (is a list/array).

</details>

<details>

<summary>Supplier Company Registration</summary>

A list of company registration details including type and number, for the supplier of the invoice.

Accessor: `supplier_company_registration`

**Subfields**

* **Number**\
  The company registration number.\
  Accessor: `number`\
  Value Type: `string`
* **Type**\
  The type of the company registration number.\
  Accessor: `type`\
  Possible Values: `VAT`, `SIRET`, `SIREN`, `NIF`, `CF`, `UID`, `STNR`, `HRA_HRB`, `TIN`, `RFC`, `BTW`, `ABN`, `UEN`, `CVR`, `ORGNRO`, `INN`, `DPH`, `NIP`, `GSTIN`, `CRN`, `KVK`, `DIC`, `TAX_ID`, `CIF`, `GST_HST_CA`, `COC`

Can have multiple values (is a list/array).

</details>

<details>

<summary>Invoice Number</summary>

The number of the invoice.

Accessor: `invoice_number`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Date</summary>

The date the invoice was issued.

Accessor: `date`\
Value Type: `date`

Has a single value.

</details>

<details>

<summary>Total Amount</summary>

The final total amount paid, including all taxes and discounts.

Accessor: `total_amount`\
Value Type: `number`

Has a single value.

</details>

<details>

<summary>Total Net</summary>

The total amount before taxes.

Accessor: `total_net`\
Value Type: `number`

Has a single value.

</details>

<details>

<summary>Total Tax</summary>

The total amount of all taxes.

Accessor: `total_tax`\
Value Type: `number`

Has a single value.

</details>

<details>

<summary>Taxes</summary>

A list of individual taxes applied, each including rate, base and amount.

Accessor: `taxes`

**Subfields**

* **Rate**\
  The tax rate of this tax as a decimal.\
  Accessor: `rate`\
  Value Type: `number`
* **Base**\
  The base amount on which this tax is computed.\
  Accessor: `base`\
  Value Type: `number`
* **Amount**\
  The computed tax amount for this tax.\
  Accessor: `amount`\
  Value Type: `number`

Can have multiple values (is a list/array).

</details>

<details>

<summary>Line Items</summary>

A list of line items in the invoice.

Accessor: `line_items`

**Subfields**

* **Description**\
  A description of the item or service.\
  Accessor: `description`\
  Value Type: `string`
* **Quantity**\
  The quantity of the item or service.\
  Accessor: `quantity`\
  Value Type: `number`
* **Unit Price**\
  The price per unit of the item or service.\
  Accessor: `unit_price`\
  Value Type: `number`
* **Total Price**\
  The total price for the line item: quantity \* unit price.\
  Accessor: `total_price`\
  Value Type: `number`
* **Product Code**\
  The product code of the item.\
  Accessor: `product_code`\
  Value Type: `string`
* **Tax Amount**\
  The tax amount of the item.\
  Accessor: `tax_amount`\
  Value Type: `number`
* **Tax Rate**\
  The tax rate of the item.\
  Accessor: `tax_rate`\
  Value Type: `number`
* **Unit Measure**\
  The unit of measure of the item.\
  Accessor: `unit_measure`\
  Value Type: `string`

Can have multiple values (is a list/array).

</details>

<details>

<summary>Document Type</summary>

Document type of the financial document.

Accessor: `document_type`\
Possible Values: `invoice`, `payslip`, `quote`, `purchase_order`, `statement`, `receipt`, `credit_note`, `other_financial`

Has a single value.

</details>

<details>

<summary>Locale</summary>

The locale contains the language, country and currency of the invoice.

Accessor: `locale`

**Subfields**

* **Language**\
  The language of the invoice, ISO 639-1 language code.\
  Accessor: `language`\
  Value Type: `string`
* **Country**\
  The country of the invoice, ISO 3166-1 alpha-2.\
  Accessor: `country`\
  Value Type: `string`
* **Currency**\
  The currency in which is issued the invoice, ISO 4217.\
  Accessor: `currency`\
  Value Type: `string`

Has a single value.

</details>

<details>

<summary>Customer Name</summary>

The name of the customer of the invoice.

Accessor: `customer_name`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Customer Address</summary>

The full address of the customer and the breakdown of the address in components.

Accessor: `customer_address`

**Subfields**

* **Address**\
  The full raw address of the customer as written in the document.\
  Accessor: `address`\
  Value Type: `string`
* **Street number**\
  The street number of the address.\
  Accessor: `street_number`\
  Value Type: `string`
* **Street Name**\
  The street name of the address.\
  Accessor: `street_name`\
  Value Type: `string`
* **PO Box**\
  PO Box number, if there is one in the address.\
  Accessor: `po_box`\
  Value Type: `string`
* **Address Complement**\
  Address complement: floor, building, suite, ...\
  Accessor: `address_complement`\
  Value Type: `string`
* **City**\
  The city of the address.\
  Accessor: `city`\
  Value Type: `string`
* **Postal Code**\
  The postal code of the address.\
  Accessor: `postal_code`\
  Value Type: `string`
* **State**\
  The state / region / land of the address, if there is any.\
  Accessor: `state`\
  Value Type: `string`
* **Country**\
  The country of the address.\
  Accessor: `country`\
  Value Type: `string`

Has a single value.

</details>

<details>

<summary>Shipping Address</summary>

The full shipping address and the breakdown of the address in components.

Accessor: `shipping_address`

**Subfields**

* **Address**\
  The full raw shipping address as written in the document.\
  Accessor: `address`\
  Value Type: `string`
* **Street number**\
  The street number of the address.\
  Accessor: `street_number`\
  Value Type: `string`
* **Street Name**\
  The street name of the address.\
  Accessor: `street_name`\
  Value Type: `string`
* **PO Box**\
  PO Box number, if there is one in the address.\
  Accessor: `po_box`\
  Value Type: `string`
* **Address Complement**\
  Address complement: floor, building, suite, ...\
  Accessor: `address_complement`\
  Value Type: `string`
* **City**\
  The city of the address.\
  Accessor: `city`\
  Value Type: `string`
* **Postal Code**\
  The postal code of the address.\
  Accessor: `postal_code`\
  Value Type: `string`
* **State**\
  The state / region / land of the address, if there is any.\
  Accessor: `state`\
  Value Type: `string`
* **Country**\
  The country of the address.\
  Accessor: `country`\
  Value Type: `string`

Has a single value.

</details>

<details>

<summary>Billing Address</summary>

The full billing address and the breakdown of the address in components.

Accessor: `billing_address`

**Subfields**

* **Address**\
  The full full raw billing address as written in the document.\
  Accessor: `address`\
  Value Type: `string`
* **Street number**\
  The street number of the address.\
  Accessor: `street_number`\
  Value Type: `string`
* **Street Name**\
  The street name of the address.\
  Accessor: `street_name`\
  Value Type: `string`
* **PO Box**\
  PO Box number, if there is one in the address.\
  Accessor: `po_box`\
  Value Type: `string`
* **Address Complement**\
  Address complement: floor, building, suite, ...\
  Accessor: `address_complement`\
  Value Type: `string`
* **City**\
  The city of the address.\
  Accessor: `city`\
  Value Type: `string`
* **Postal Code**\
  The postal code of the address.\
  Accessor: `postal_code`\
  Value Type: `string`
* **State**\
  The state / region / land of the address, if there is any.\
  Accessor: `state`\
  Value Type: `string`
* **Country**\
  The country of the address.\
  Accessor: `country`\
  Value Type: `string`

Has a single value.

</details>

<details>

<summary>Supplier Address</summary>

The full address of the supplier and the breakdown of the address in components.

Accessor: `supplier_address`

**Subfields**

* **Address**\
  The full raw supplier address as written in the document.\
  Accessor: `address`\
  Value Type: `string`
* **Street number**\
  The street number of the address.\
  Accessor: `street_number`\
  Value Type: `string`
* **Street Name**\
  The street name of the address.\
  Accessor: `street_name`\
  Value Type: `string`
* **PO Box**\
  PO Box number, if there is one in the address.\
  Accessor: `po_box`\
  Value Type: `string`
* **Address Complement**\
  Address complement: floor, building, suite, ...\
  Accessor: `address_complement`\
  Value Type: `string`
* **City**\
  The city of the address.\
  Accessor: `city`\
  Value Type: `string`
* **Postal Code**\
  The postal code of the address.\
  Accessor: `postal_code`\
  Value Type: `string`
* **State**\
  The state / region / land of the address, if there is any.\
  Accessor: `state`\
  Value Type: `string`
* **Country**\
  The country of the address.\
  Accessor: `country`\
  Value Type: `string`

Has a single value.

</details>

<details>

<summary>Due Date</summary>

The date on which the invoice is due.

Accessor: `due_date`\
Value Type: `date`

Has a single value.

</details>

<details>

<summary>PO Number</summary>

The purchase order number.

Accessor: `po_number`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Reference Numbers</summary>

List of all reference numbers on the invoice, including the purchase order number.

Accessor: `reference_numbers`\
Value Type: `string`

Can have multiple values (is a list/array).

</details>

<details>

<summary>Payment Date</summary>

The date on which the payment is due / was full-filled.

Accessor: `payment_date`\
Value Type: `date`

Has a single value.

</details>

<details>

<summary>Supplier Payment Details</summary>

List of payment details associated to the supplier of the invoice.

Accessor: `supplier_payment_details`

**Subfields**

* **IBAN**\
  The International Bank Account Number (IBAN).\
  Accessor: `iban`\
  Value Type: `string`
* **SWIFT**\
  The bank s SWIFT Business Identifier Code (BIC).\
  Accessor: `swift`\
  Value Type: `string`
* **Account Number**\
  The account number.\
  Accessor: `account_number`\
  Value Type: `string`
* **Routing Number**\
  The routing number.\
  Accessor: `routing_number`\
  Value Type: `string`

Can have multiple values (is a list/array).

</details>

<details>

<summary>Supplier Website</summary>

The website URL of the supplier or merchant.

Accessor: `supplier_website`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Supplier Email</summary>

The email address of the supplier or merchant.

Accessor: `supplier_email`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Customer ID</summary>

The customer account number or identifier from the supplier.

Accessor: `customer_id`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Purchase Category</summary>

The category of the receipt.

Accessor: `category`\
Possible Values: `food`, `gasoline`, `parking`, `toll`, `accommodation`, `transport`, `telecom`, `software`, `shopping`, `energy`, `miscellaneous`

Has a single value.

</details>

<details>

<summary>Purchase Subcategory</summary>

The purchase subcategory of the receipt.

Accessor: `subcategory`\
Possible Values: `restaurant`, `delivery`, `train`, `public`, `taxi`, `car_rental`, `plane`, `micromobility`, `office_supplies`, `electronics`, `cultural`, `groceries`, `other`

Has a single value.

</details>

<details>

<summary>Receipt Number</summary>

The receipt number or identifier only if document is a receipt.

Accessor: `receipt_number`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Document Number</summary>

The document number or identifier (invoice number or receipt number).

Accessor: `document_number`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Tip and Gratuity</summary>

The total amount of tip and gratuity

Accessor: `tip`\
Value Type: `number`

Has a single value.

</details>

<details>

<summary>Purchase Time</summary>

The time the purchase was made (only for receipts).

Accessor: `time`\
Value Type: `string`

Has a single value.

</details>


# International ID Card

Mindee enables the automatic extraction of structured identity data from ID cards, a critical component for use cases like identity verification, KYC, and onboarding.

Watch this short demo to see how quickly you can set up and test the French ID Card model with Mindee:

{% @supademo/embed url="<https://app.supademo.com/demo/cmf59btop30t539ozuva4neqi>" demoId="cmf59btop30t539ozuva4neqi" %}

## Why use Mindee for ID cards?

ID cards differ vastly across countries in design, language, and layout. Mindee abstracts that complexity by offering:

* Pre-trained models ready for different document types
* Plug-and-play integration—no template design or training required
* High accuracy and speed, even with variations in card format, font, or photo quality

## What can be extracted from ID cards?

Mindee’s ID card API returns structured fields found on most ID documents:

| Field                  | Description                        |
| ---------------------- | ---------------------------------- |
| Document Type          | Type of ID (like national ID card) |
| Document Number        | Unique Identifier on the card      |
| Given Name(s), Surname | Personal name data                 |
| Birth Date             | YYYY-MM-DD format                  |
| Birth Place            | City or region of birth            |
| Issurance Date         | Card issue date                    |
| Expiry Date            | Expiration date                    |
| Gender                 | M/F/other options                  |
| Nationality            | Issuing country code               |
| MRZ Lines              | If present, machine-readable data  |
| Address                | (when applicable)                  |
| Additional metada      | Depends on region/document         |
| Signature              | Signature of the holder            |
| ID photo               | Photo of the holder                |

## Example: French National ID Card (Carte Nationale d’Identité – France)

### Why Mindee works for French ID Cards

Mindee’s ID Card France OCR model handles all French ID variations. It’s optimized to read both MRZ and non‑MRZ zones, ensuring high accuracy across all formats.

Mindee can extract specific French ID card fields like:

* **Card Access Number (CAN)** – a unique numeric code specific to French ID cards; not labeled on the document itself
* **Issuing Authority** – indicates the specific French administrative body (mairie, préfecture, etc.) that issued the card
* **Document Type** – classification indicating which version of the ID card it is (NEW - the issued version from 2021 or OLD - the issued version from 1988 - 2021)
* **Document Sides** – indicates whether the processed image is the front or back of the ID
* **Image Orientation** – rotation metadata indicating how the image was oriented during extraction

If you want to try and do a live test, you can use this French ID example:

<figure><img src="/files/0os8pKRVYpD0OjMdnTLI" alt="" width="563"><figcaption></figcaption></figure>

### Two Ways to Start

### 1. Choose "International ID" in the Catalog (Recommended)

* Click on "Create your document AI model" in your dashboard, then select **"**&#x49;nternational I&#x44;**".**
* The International ID model template comes pre-configured with standard [#international-id-fields](#international-id-fields "mention").
* Once your Invoice model is created, you can immediately [test](/models/live-test) with your own invoices.
* Optionally, you can adjust the model's [Data Schema](/extraction-models/data-schema) if you need to modify fields.

#### 1. Pick the ID Card Model from the Catalog (Recommended)

* In your Mindee dashboard, go to the **Document Catalog** and choose the **“International ID”** model.
* Once selected, the platform will generate a prefilled schema with standard identity fields. You can accept that schema as is or adjust it in the **Data Schema** section to fine-tune what’s extracted.
* After the model is set up, you can immediately test it with your own documents.

#### 2. Build an ID Card Model from Scratch

* In the dialog with Mindee’s AI assistant, describe what the document is (e.g., “a French national ID card”) and specify the fields you want to extract (such as surname, date of birth, issuing authority).
* The AI will propose an initial schema—review it and adjust fields or mapping in the **Data Schema** tab until it fits your requirements.
* Once finalized, your custom model is ready for live testing.

## Document format support

The API handles both PDF and common image formats (JPG, PNG), including smartphone photos. It applies reliably to front/back sides and both old and new French ID layouts.

## International ID Fields

Documentation for all fields present in the data schema.

Field *accessors* are used as the keys for accessing the values in the returned data. On the data schema interface, this is the "Field Name".

Field *value types* indicate how the value is returned by the API. On the data schema interface, this is the "Field Type".

<details>

<summary>Given Names</summary>

The given names of the person.

Accessor: `given_names`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Surnames</summary>

The surnames of the person.

Accessor: `surnames`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Date of Birth</summary>

The date of birth of the person.

Accessor: `date_of_birth`\
Value Type: `date`

Has a single value.

</details>

<details>

<summary>Place of Birth</summary>

The place of birth of the person.

Accessor: `place_of_birth`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Nationality</summary>

The nationality of the person.

Accessor: `nationality`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Sex</summary>

The sex of the person.

Accessor: `sex`\
Possible Values: `M`, `F`, `Other`

Has a single value.

</details>

<details>

<summary>Document Number</summary>

The document number of the ID.

Accessor: `document_number`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Date of Issue</summary>

The date when the ID was issued.

Accessor: `date_of_issue`\
Value Type: `date`

Has a single value.

</details>

<details>

<summary>Date of Expiry</summary>

The date when the ID expires.

Accessor: `date_of_expiry`\
Value Type: `date`

Has a single value.

</details>

<details>

<summary>Authority</summary>

The authority that issued the ID.

Accessor: `authority`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Address</summary>

The address of the person.

Accessor: `address`

**Subfields**

* **Street**\
  The street address.\
  Accessor: `street`\
  Value Type: `string`
* **City**\
  The city.\
  Accessor: `city`\
  Value Type: `string`
* **State**\
  The state or province.\
  Accessor: `state`\
  Value Type: `string`
* **Postal Code**\
  The postal code.\
  Accessor: `postal_code`\
  Value Type: `string`
* **Country**\
  The country.\
  Accessor: `country`\
  Value Type: `string`

Has a single value.

</details>

<details>

<summary>Document Type</summary>

The type of ID document.

Accessor: `document_type`\
Possible Values: `Passport`, `National ID`, `Driver License`, `Residence Permit`

Has a single value.

</details>

<details>

<summary>MRZ</summary>

The mrz, the Machine Readable Zone.

Accessor: `mrz`

**Subfields**

* **Line 1**\
  The first line of the mrz.\
  Accessor: `line_1`\
  Value Type: `string`
* **Line 2**\
  The second line of the mrz.\
  Accessor: `line_2`\
  Value Type: `string`
* **Line 3**\
  The third line of the mrz.\
  Accessor: `line_3`\
  Value Type: `string`

Has a single value.

</details>


# Passport

Whether you are building a Know Your Customer (KYC) process, onboarding flow, or document verification pipeline, Mindee helps you parse structured passport data with accuracy.

Take a look at our demo for the Indian Passport model:

{% @supademo/embed url="<https://app.supademo.com/demo/cmeqvjfgw3holv9kqocym86w9>" demoId="cmeqvjfgw3holv9kqocym86w9" %}

## Why Use Mindee for Passports?

Passports vary in format depending on the country, language, and issuance authority. Mindee simplifies extraction by letting you:

* Describe which fields matter to you
* Upload sample documents to refine extraction
* Get structured outputs without training models yourself

## Two Ways to Start Building your Passport Model

### 1. Choose "Passport" in the Catalog (Recommended)

* Click on "Create your document AI model" in your dashboard, then select **"Passport".**
* The Passport model template comes pre-configured with standard [#passport-fields](#passport-fields "mention").
* Once your Invoice model is created, you can immediately [test](/models/live-test) with your own invoices.
* Optionally, you can adjust the model's [Data Schema](/extraction-models/data-schema) if you need to modify fields.

### 2. Build a Passport Model from Scratch

* Ask the specifications directly to the AI assistant (e.g. "*I want a model that extracts the following fields from passports: surname, date of birth and MRZ*").
* Additionally, you can upload a sample passport if needed for clarification.
* Mindee builds you a tailored parser in seconds.

If you want to try and do a live test, here is a sample for Indian passport:

<figure><img src="/files/z0j9yKFayrEOe9h3Nn3z" alt="" width="563"><figcaption></figcaption></figure>

<figure><img src="/files/zQOqlsoenyup9DQsGO8b" alt="" width="563"><figcaption></figcaption></figure>

### Indian Passport Example

Indian passports include additional region-specific fields. If you're processing Indian passports, our technology is also able to extract, for example, those fields:

| Field                       | Description                                  |
| --------------------------- | -------------------------------------------- |
| Name of Legal Guardian      | Often the father’s name or guardian’s name   |
| Name of Spouse              | Appears if applicable                        |
| Name of Mother              | Optional field                               |
| Old Passport Number         | Prior document if applicable                 |
| Old Passport Date of Issue  | Issue date of prior passport                 |
| Old Passport Place of Issue | Location where the prior passport was issued |
| File Number                 | Government-issued file reference             |
| Address Line 1–3            | Full residential address                     |

{% hint style="info" %}
You can just upload an Indian passport and ask to extract the fields present.
{% endhint %}

## Document Format

* We support both **single-page and multi-page PDFs/images of passports.**
* You can add the fields of any page in the data schema as they are all supported by Mindee.

## Passport Fields

Documentation for all fields present in the data schema.

Field *accessors* are used as the keys for accessing the values in the returned data. On the data schema interface, this is the "Field Name".

Field *value types* indicate how the value is returned by the API. On the data schema interface, this is the "Field Type".

<details>

<summary>Given Names</summary>

The given names (first names) of the passport holder.

Accessor: `given_names`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Surnames</summary>

The surnames (last names) of the passport holder.

Accessor: `surnames`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Date of Birth</summary>

The date of birth of the passport holder.

Accessor: `date_of_birth`\
Value Type: `date`

Has a single value.

</details>

<details>

<summary>Place of Birth</summary>

The place of birth of the passport holder.

Accessor: `place_of_birth`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Passport Number</summary>

The passport number.

Accessor: `passport_number`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Issuing Country</summary>

The country that issued the passport.

Accessor: `issuing_country`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Nationality</summary>

The nationality of the passport holder.

Accessor: `nationality`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Date of Issue</summary>

The date the passport was issued.

Accessor: `date_of_issue`\
Value Type: `date`

Has a single value.

</details>

<details>

<summary>Date of Expiry</summary>

The date the passport expires.

Accessor: `date_of_expiry`\
Value Type: `date`

Has a single value.

</details>

<details>

<summary>Sex</summary>

The sex of the passport holder.

Accessor: `sex`\
Possible Values: `Male`, `Female`, `Other`

Has a single value.

</details>

<details>

<summary>MRZ Line 1</summary>

The first line of the Machine Readable Zone (MRZ).

Accessor: `mrz_line_1`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>MRZ Line 2</summary>

The second line of the Machine Readable Zone (MRZ).

Accessor: `mrz_line_2`\
Value Type: `string`

Has a single value.

</details>


# Driver's License

Automatically parse driver licenses and extract structured driver data using the Driver License model available in the Catalog.

Documentation for the data schema of the Driver's License model template.

This template can extract data from driver license documents issued by any country.

Some fields apply only to some countries or jurisdictions, if you have no need for these simply remove them. Conversely, you may need to add extra fields in order to support your specific business rules.

## Driver's License Fields

Documentation for all fields present in the data schema.

Field *accessors* are used as the keys for accessing the values in the returned data. On the data schema interface, this is the "Field Name".

Field *value types* indicate how the value is returned by the API. On the data schema interface, this is the "Field Type".

<details>

<summary>First Name</summary>

The given name(s) or first name(s) of the person.

Accessor: `first_name`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Last Name</summary>

The surnames or last names of the person.

Accessor: `last_name`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Date of Birth</summary>

The date of birth of the person.

Accessor: `date_of_birth`\
Value Type: `date`

Has a single value.

</details>

<details>

<summary>Place of Birth</summary>

The place of birth of the person.

Accessor: `place_of_birth`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Nationality</summary>

The nationality of the person.

Accessor: `nationality`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Sex</summary>

The sex of the person.

Accessor: `sex`\
Possible Values: `M`, `F`, `Other`

Has a single value.

</details>

<details>

<summary>Document Id</summary>

The document number or the ID number of the document.

Accessor: `document_id`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Issued Date</summary>

The date when the document was issued.

Accessor: `issued_date`\
Value Type: `date`

Has a single value.

</details>

<details>

<summary>Expiry Date</summary>

The date when the Document expires.

Accessor: `expiry_date`\
Value Type: `date`

Has a single value.

</details>

<details>

<summary>Country Code</summary>

Country code extracted as a string.

Accessor: `country_code`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Issuing Authority</summary>

The authority that issued the document.

Accessor: `issuing_authority`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Address</summary>

The address of the person.

Accessor: `address`

**Subfields**

* **Street**\
  The street address.\
  Accessor: `street`\
  Value Type: `string`
* **City**\
  The city.\
  Accessor: `city`\
  Value Type: `string`
* **State**\
  The state or province.\
  Accessor: `state`\
  Value Type: `string`
* **Postal Code**\
  The postal code.\
  Accessor: `postal_code`\
  Value Type: `string`
* **Country**\
  The country.\
  Accessor: `country`\
  Value Type: `string`

Has a single value.

</details>

<details>

<summary>MRZ</summary>

Machine-readable license number

Accessor: `mrz`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Driver License Category</summary>

EU driver license holders categories

Accessor: `category`\
Value Type: `string`

Has a single value.

</details>


# Boarding Pass

With Mindee, you can extract all the information needed for a passenger to access a flight: traveler identity, flight details, departure and arrival airports, and boarding information.

Here is how to build your custom Boarding Pass model with Mindee in just a few minutes:

{% @supademo/embed url="<https://app.supademo.com/demo/cmfl5e3l1erva39ozltfkx1ta>" demoId="cmfl5e3l1erva39ozltfkx1ta" %}

## Why use Mindee for boarding passes?

Boarding passes differ depending on airline, format (paper or mobile), and layout. Instead of handling these variations yourself, you can simply tell the AI Agent which fields you want to extract, and it will design the model for you.

Common use cases:

* Travel and mobility apps verifying passenger details
* Loyalty program enrollment and automation
* Expense reporting and travel management

### What can be extracted from boarding passes?

Typical fields you may want to capture include:

| Field                       | Description                                   |
| --------------------------- | --------------------------------------------- |
| Passenger Name              | Full name of the traveler                     |
| Flight Number               | Airline code and flight number                |
| Departure Date              | Date of departure in YYYY-MM-DD               |
| Boarding Time               | Boarding time                                 |
| Departure Airport           | Name and IATA code of the origin airport      |
| Destination Airport         | Name and IATA code of the destination airport |
| Bardcode (Object Detection) | Image of the barcode                          |

Depending on your needs, you can also add fields such as **Seat Number**, **Gate**, or **Booking Reference (PNR)**.

Here is a boarding pass sample if you want to try a live test for your model:

<figure><img src="/files/pTCDdfl7Coj1pqFxB1mO" alt="a fake boarding pass"><figcaption></figcaption></figure>

## How to Create your Boarding Pass Model

Since there is no boarding pass Model Template in the Catalog, you will need to build your own custom model using the AI Agent:

1. Describe what you want to the AI assistant (e.g. *“I want to capture passenger name, flight number, departure date, and airports from boarding passes. Also add an object detection field for the barcode zone.”*)
2. Optionally upload a sample boarding pass to give more context.
3. The Agent will generate a model for you and propose a set of fields.
4. You can refine the schema by asking for more or fewer fields until it matches your requirements.
5. Once ready, you can test the model live with your documents.

## Document format support

The custom model accepts PDF and common image formats (JPG, PNG). It works with digital boarding passes as well as photos or scans of printed passes. See full [list of accepted files](https://docs.mindee.com/integrations/technical-limitations#accepted-files).


# Resume

Automatically parse résumés/CVs and extract structured candidate data using the pre-trained Resume model available in the Catalog.

Documentation for the data schema of the Bill of Lading model template.

## Resume Fields

Documentation for all fields present in the data schema.

Field *accessors* are used as the keys for accessing the values in the returned data. On the data schema interface, this is the "Field Name".

Field *value types* indicate how the value is returned by the API. On the data schema interface, this is the "Field Type".

<details>

<summary>Name</summary>

The full name of the resume owner

Accessor: `name`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Address</summary>

The address of the resume owner

Accessor: `address`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Phone Number</summary>

The phone number of the resume owner

Accessor: `phone_number`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Email</summary>

The email address of the resume owner

Accessor: `email`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>LinkedIn Profile</summary>

The LinkedIn profile URL of the resume owner

Accessor: `linkedin_profile`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Education</summary>

List of educational experiences

Accessor: `education`

**Subfields**

* **School Name**\
  The name of the school\
  Accessor: `school_name`\
  Value Type: `string`
* **Degree**\
  The degree obtained\
  Accessor: `degree`\
  Value Type: `string`
* **Dates Attended**\
  The dates the person attended the school\
  Accessor: `dates_attended`\
  Value Type: `string`
* **GPA**\
  The GPA obtained\
  Accessor: `gpa`\
  Value Type: `number`
* **Relevant Coursework**\
  Relevant coursework taken\
  Accessor: `relevant_coursework`\
  Value Type: `string`

Can have multiple values (is a list/array).

</details>

<details>

<summary>Experience</summary>

List of work experiences

Accessor: `experience`

**Subfields**

* **Company Name**\
  The name of the company\
  Accessor: `company_name`\
  Value Type: `string`
* **Job Title**\
  The job title held\
  Accessor: `job_title`\
  Value Type: `string`
* **Dates Employed**\
  The dates the person was employed\
  Accessor: `dates_employed`\
  Value Type: `string`
* **Responsibilities**\
  List of responsibilities\
  Accessor: `responsibilities`\
  Value Type: `string`

Can have multiple values (is a list/array).

</details>

<details>

<summary>Skills</summary>

List of skills

Accessor: `skills`\
Value Type: `string`

Can have multiple values (is a list/array).

</details>

<details>

<summary>Languages</summary>

List of languages and proficiency

Accessor: `languages`

**Subfields**

* **Language**\
  The language\
  Accessor: `language`\
  Value Type: `string`
* **Proficiency**\
  The proficiency level\
  Accessor: `proficiency`\
  Value Type: `string`

Can have multiple values (is a list/array).

</details>

<details>

<summary>Projects</summary>

List of projects

Accessor: `projects`

**Subfields**

* **Project Name**\
  The name of the project\
  Accessor: `project_name`\
  Value Type: `string`
* **Description**\
  The description of the project\
  Accessor: `description`\
  Value Type: `string`
* **Technologies Used**\
  List of technologies used\
  Accessor: `technologies_used`\
  Value Type: `string`

Can have multiple values (is a list/array).

</details>

<details>

<summary>Awards &#x26; Certifications</summary>

List of awards and certifications

Accessor: `awards_certifications`

**Subfields**

* **Name**\
  The name of the award or certification\
  Accessor: `name`\
  Value Type: `string`
* **Date Received**\
  The date the award or certification was received\
  Accessor: `date_received`\
  Value Type: `date`

Can have multiple values (is a list/array).

</details>

<details>

<summary>Summary/Objective</summary>

The summary or objective statement

Accessor: `summary_objective`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Candidate Photo</summary>

The profile photo of the candidate.

Accessor: `candidate_photo`\
Value Type: `object_detection`

Has a single value.

</details>


# Bill of Lading

Automatically parse Bills of Lading and extract structured shipping data using the pre-trained Bill of Lading model template available in the Catalog.

Documentation for the data schema of the Bill of Lading model template.

## Bill of Lading Fields

Documentation for all fields present in the data schema.

Field *accessors* are used as the keys for accessing the values in the returned data. On the data schema interface, this is the "Field Name".

Field *value types* indicate how the value is returned by the API. On the data schema interface, this is the "Field Type".

<details>

<summary>Bill of Lading Number</summary>

A unique identifier assigned to a Bill of Lading document

Accessor: `bill_of_lading_number`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Shipper</summary>

The party responsible for shipping the goods

Accessor: `shipper`

**Subfields**

* **Name**\
  Shipper s name\
  Accessor: `name`\
  Value Type: `string`
* **Address**\
  Shipper s address\
  Accessor: `address`\
  Value Type: `string`
* **Phone**\
  Shipper s phone number\
  Accessor: `phone`\
  Value Type: `string`
* **Email**\
  Shipper s email address\
  Accessor: `email`\
  Value Type: `string`

Has a single value.

</details>

<details>

<summary>Consignee</summary>

The party to whom the goods are being shipped

Accessor: `consignee`

**Subfields**

* **Name**\
  Consignee s name\
  Accessor: `name`\
  Value Type: `string`
* **Address**\
  Consignee s address\
  Accessor: `address`\
  Value Type: `string`
* **Phone**\
  Consignee s phone number\
  Accessor: `phone`\
  Value Type: `string`
* **Email**\
  Consignee s email address\
  Accessor: `email`\
  Value Type: `string`

Has a single value.

</details>

<details>

<summary>Notify Party</summary>

The party to be notified of the arrival of the goods

Accessor: `notify_party`

**Subfields**

* **Name**\
  Notify party s name\
  Accessor: `name`\
  Value Type: `string`
* **Address**\
  Notify party s address\
  Accessor: `address`\
  Value Type: `string`
* **Phone**\
  Notify party s phone number\
  Accessor: `phone`\
  Value Type: `string`
* **Email**\
  Notify party s email address\
  Accessor: `email`\
  Value Type: `string`

Has a single value.

</details>

<details>

<summary>Carrier</summary>

The shipping company responsible for transporting the goods

Accessor: `carrier`

**Subfields**

* **Name**\
  Carrier s name\
  Accessor: `name`\
  Value Type: `string`
* **Professional Number**\
  Carrier s professional number\
  Accessor: `professional_number`\
  Value Type: `string`
* **SCAC**\
  Carrier s SCAC code\
  Accessor: `scac`\
  Value Type: `string`

Has a single value.

</details>

<details>

<summary>Carrier Items</summary>

The goods being shipped

Accessor: `carrier_items`

**Subfields**

* **Description**\
  Description of the item\
  Accessor: `description`\
  Value Type: `string`
* **Quantity**\
  Quantity of the item\
  Accessor: `quantity`\
  Value Type: `number`
* **Gross Weight**\
  Gross weight of the item\
  Accessor: `gross_weight`\
  Value Type: `number`
* **Weight Unit**\
  Unit of weight\
  Accessor: `weight_unit`\
  Value Type: `string`
* **Measurement**\
  Measurement of the item\
  Accessor: `measurement`\
  Value Type: `number`
* **Measurement Unit**\
  Unit of measurement\
  Accessor: `measurement_unit`\
  Value Type: `string`

Can have multiple values (is a list/array).

</details>

<details>

<summary>Port of Loading</summary>

The port where the goods are loaded onto the vessel

Accessor: `port_of_loading`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Port of Discharge</summary>

The port where the goods are unloaded from the vessel

Accessor: `port_of_discharge`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Place of Delivery</summary>

The place where the goods are to be delivered

Accessor: `place_of_delivery`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Issue Date</summary>

The date when the bill of lading is issued.

Accessor: `issue_date`\
Value Type: `date`

Has a single value.

</details>

<details>

<summary>Departure Date</summary>

The date when the vessel departs from the port of loading

Accessor: `departure_date`\
Value Type: `date`

Has a single value.

</details>


# Bank Statement

Automatically parse bank statements and extract structured financial data using the pre-trained Bank Statement model template available in the Catalog.

Documentation for the data schema of the Bank Statement model template.

## Bank Statement Fields

Documentation for all fields present in the data schema.

Field *accessors* are used as the keys for accessing the values in the returned data. On the data schema interface, this is the "Field Name".

Field *value types* indicate how the value is returned by the API. On the data schema interface, this is the "Field Type".

<details>

<summary>Account Holder Name</summary>

List of the account holder s names

Accessor: `account_holder_names`\
Value Type: `string`

Can have multiple values (is a list/array).

</details>

<details>

<summary>Account Number</summary>

Account number

Accessor: `account_number`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Statement Period Start Date</summary>

Start date of the statement period

Accessor: `statement_period_start_date`\
Value Type: `date`

Has a single value.

</details>

<details>

<summary>Statement Period End Date</summary>

End date of the statement period

Accessor: `statement_period_end_date`\
Value Type: `date`

Has a single value.

</details>

<details>

<summary>Statement Date</summary>

Date the statement was issued

Accessor: `statement_date`\
Value Type: `date`

Has a single value.

</details>

<details>

<summary>Beginning Balance</summary>

Beginning balance of the statement period

Accessor: `beginning_balance`\
Value Type: `number`

Has a single value.

</details>

<details>

<summary>Ending Balance</summary>

Ending balance of the statement period

Accessor: `ending_balance`\
Value Type: `number`

Has a single value.

</details>

<details>

<summary>List of Transactions</summary>

List of transactions

Accessor: `list_of_transactions`

**Subfields**

* **Date**\
  Date of the transaction\
  Accessor: `date`\
  Value Type: `date`
* **Description**\
  Description of the transaction\
  Accessor: `description`\
  Value Type: `string`
* **Amount**\
  Amount of the transaction\
  Accessor: `amount`\
  Value Type: `number`

Can have multiple values (is a list/array).

</details>

<details>

<summary>Bank Name</summary>

Name of the bank

Accessor: `bank_name`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Account Type</summary>

Type of account

Accessor: `account_type`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Currency</summary>

Currency of the account

Accessor: `currency`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Total Credits</summary>

Total credits for the statement period

Accessor: `total_credits`\
Value Type: `number`

Has a single value.

</details>

<details>

<summary>Total Debits</summary>

Total debits for the statement period

Accessor: `total_debits`\
Value Type: `number`

Has a single value.

</details>

<details>

<summary>Bank Address</summary>

Address of the bank

Accessor: `bank_address`

**Subfields**

* **Address**\
  The full raw address of the bank as written in the document.\
  Accessor: `address`\
  Value Type: `string`
* **Street**\
  The street address.\
  Accessor: `street`\
  Value Type: `string`
* **City**\
  The city.\
  Accessor: `city`\
  Value Type: `string`
* **State**\
  The state or province.\
  Accessor: `state`\
  Value Type: `string`
* **Postal Code**\
  The postal code.\
  Accessor: `postal_code`\
  Value Type: `string`
* **Country**\
  The country.\
  Accessor: `country`\
  Value Type: `string`

Has a single value.

</details>

<details>

<summary>Branch Code</summary>

Branch code of the bank

Accessor: `branch_code`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Daily Balances</summary>

List of daily balances.

Accessor: `daily_balances`

**Subfields**

* **Date**\
  The date of the daily balance item, returned as an ISO formatted string (yyyy-mm-dd).\
  Accessor: `date`\
  Value Type: `date`
* **Balance Amount**\
  Signed number representing the value of the balance.\
  Accessor: `balance_amount`\
  Value Type: `number`

Can have multiple values (is a list/array).

</details>

<details>

<summary>Account Holder Address</summary>

Address of account holder.

Accessor: `account_holder_address`

**Subfields**

* **Street**\
  The street address.\
  Accessor: `street`\
  Value Type: `string`
* **City**\
  The city.\
  Accessor: `city`\
  Value Type: `string`
* **State**\
  The state or province.\
  Accessor: `state`\
  Value Type: `string`
* **Postal Code**\
  The postal code.\
  Accessor: `postal_code`\
  Value Type: `string`
* **Country**\
  The country.\
  Accessor: `country`\
  Value Type: `string`

Has a single value.

</details>


# European Vehicle Registration

Documentation for the data schema of European Vehicle Registration model.

Extract data from Vehicle Registration Certificates issued in all EU countries.

Supported document formats include, but are not limited to:

* DEU - *Zulassungsbescheinigung Teil I* / *II*
* ESP - *Permiso de Circulación*
* FRA - *Certificat d'immatriculation* (*Carte grise*)
* ITA - *Carta di Circolazione* or *Documento Unico*

## European Vehicle Registration Fields

Documentation for all fields present in the data schema.

Field *accessors* are used as the keys for accessing the values in the returned data. On the data schema interface, this is the "Field Name".

Field *value types* indicate how the value is returned by the API. On the data schema interface, this is the "Field Type".

<details>

<summary>document_number</summary>

The unique identifier for the car technical passport document.

Accessor: `document_number`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>a</summary>

The registration number of the vehicle.

Accessor: `a`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>b</summary>

The date when the vehicle was registered.

Accessor: `b`\
Value Type: `date`

Has a single value.

</details>

<details>

<summary>c1</summary>

The name of the vehicle owner.

Accessor: `c1`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>c2</summary>

The name of the holder of the vehicle.

Accessor: `c2`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>c3</summary>

The address of the vehicle owner.

Accessor: `c3`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>c4</summary>

A declaration stating the ownership of the vehicle.

Accessor: `c4`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>c4_1</summary>

The name of the previous owner of the vehicle.

Accessor: `c4_1`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>d1</summary>

The manufacturer of the vehicle.

Accessor: `d1`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>d2</summary>

The type of the vehicle.

Accessor: `d2`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>d2_1</summary>

The model of the vehicle.

Accessor: `d2_1`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>d3</summary>

The variant of the vehicle model.

Accessor: `d3`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>e</summary>

The Vehicle Identification Number `VIN`.

Accessor: `e`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>f1</summary>

The gross weight of the vehicle in kilograms.

Accessor: `f1`\
Value Type: `number`

Has a single value.

</details>

<details>

<summary>f2</summary>

The unladen weight of the vehicle in kilograms.

Accessor: `f2`\
Value Type: `number`

Has a single value.

</details>

<details>

<summary>f3</summary>

The payload capacity of the vehicle in kilograms.

Accessor: `f3`\
Value Type: `number`

Has a single value.

</details>

<details>

<summary>g</summary>

The engine capacity of the vehicle in cubic centimeters.

Accessor: `g`\
Value Type: `number`

Has a single value.

</details>

<details>

<summary>g1</summary>

The net engine capacity of the vehicle in cubic centimeters.

Accessor: `g1`\
Value Type: `number`

Has a single value.

</details>

<details>

<summary>h</summary>

The validity period of the document.

Accessor: `h`\
Value Type: `date`

Has a single value.

</details>

<details>

<summary>i</summary>

Date of delivery of the present document.

Accessor: `i`\
Value Type: `date`

Has a single value.

</details>

<details>

<summary>j</summary>

The category of the vehicle as per regulations

Accessor: `j`\
Possible Values: `M1`, `N1`, `N2`, `N3`, `L`, `O`

Has a single value.

</details>

<details>

<summary>j1</summary>

The subcategory of the vehicle.

Accessor: `j1`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>j2</summary>

The approval code of the vehicle.

Accessor: `j2`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>j3</summary>

The emission class of the vehicle.

Accessor: `j3`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>k</summary>

The type approval number of the vehicle.

Accessor: `k`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>p1</summary>

The power of the vehicle in kilowatts.

Accessor: `p1`\
Value Type: `number`

Has a single value.

</details>

<details>

<summary>p2</summary>

The revolutions per minute `RPM` of the vehicle.

Accessor: `p2`\
Value Type: `number`

Has a single value.

</details>

<details>

<summary>p3</summary>

The type of fuel used by the vehicle.

Accessor: `p3`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>p5</summary>

The engine code of the vehicle.

Accessor: `p5`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>p6</summary>

The maximum speed of the vehicle in kilometers per hour.

Accessor: `p6`\
Value Type: `number`

Has a single value.

</details>

<details>

<summary>s1</summary>

The noise level of the vehicle in decibels.

Accessor: `s1`\
Value Type: `number`

Has a single value.

</details>

<details>

<summary>x1</summary>

Additional information about the vehicle.

Accessor: `x1`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>y1</summary>

The CO2 emissions of the vehicle in grams per kilometer.

Accessor: `y1`\
Value Type: `number`

Has a single value.

</details>

<details>

<summary>y2</summary>

The particulate emissions of the vehicle in grams per kilometer.

Accessor: `y2`\
Value Type: `number`

Has a single value.

</details>

<details>

<summary>y3</summary>

The hydrocarbon emissions of the vehicle in grams per kilometer.

Accessor: `y3`\
Value Type: `number`

Has a single value.

</details>

<details>

<summary>y4</summary>

The nitrogen oxide emissions of the vehicle in grams per kilometer.

Accessor: `y4`\
Value Type: `number`

Has a single value.

</details>

<details>

<summary>y5</summary>

The fuel consumption of the vehicle in liters per 100 kilometers.

Accessor: `y5`\
Value Type: `number`

Has a single value.

</details>

<details>

<summary>y6</summary>

The range of the vehicle in kilometers.

Accessor: `y6`\
Value Type: `number`

Has a single value.

</details>

<details>

<summary>z1</summary>

A reserved field for future use.

Accessor: `z1`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>z2</summary>

A reserved field for future use.

Accessor: `z2`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>z3</summary>

A reserved field for future use.

Accessor: `z3`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>z4</summary>

A reserved field for future use.

Accessor: `z4`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>z5</summary>

A reserved field for future use.

Accessor: `z5`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>mrz1</summary>

The first line of the machine-readable zone.

Accessor: `mrz1`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>mrz2</summary>

The second line of the machine-readable zone.

Accessor: `mrz2`\
Value Type: `string`

Has a single value.

</details>


# Nutrition Facts

Automatically parse food labels and extract structured nutritional data using the pre-trained Nutrition Facts model available in the Catalog.

Documentation for the data schema of the Nutrition Facts model template.

## Nutrition Facts Fields

Documentation for all fields present in the data schema.

Field *accessors* are used as the keys for accessing the values in the returned data. On the data schema interface, this is the "Field Name".

Field *value types* indicate how the value is returned by the API. On the data schema interface, this is the "Field Type".

<details>

<summary>Serving per Box</summary>

The number of servings in each box of the product

Accessor: `serving_per_box`\
Value Type: `number`

Has a single value.

</details>

<details>

<summary>Serving Size</summary>

The size of a single serving of the product

Accessor: `serving_size`

**Subfields**

* **Amount**\
  The amount of a single serving\
  Accessor: `amount`\
  Value Type: `number`
* **Unit**\
  The unit for the amount of a single serving\
  Accessor: `unit`\
  Value Type: `string`

Has a single value.

</details>

<details>

<summary>Calories</summary>

The amount of calories in the product

Accessor: `calories`

**Subfields**

* **Per 100g**\
  The amount of calories per 100g of the product\
  Accessor: `per_100g`\
  Value Type: `number`
* **Per Serving**\
  The amount of calories per serving of the product.\
  Accessor: `per_serving`\
  Value Type: `number`
* **Daily Value**\
  DVs are the recommended amounts of calories to consume or not to exceed each day\
  Accessor: `daily_value`\
  Value Type: `number`

Has a single value.

</details>

<details>

<summary>Total Fat</summary>

The total amount of fat in the product

Accessor: `total_fat`

**Subfields**

* **Per 100g**\
  The amount of total fat per 100g of the product.\
  Accessor: `per_100g`\
  Value Type: `number`
* **Per Serving**\
  The amount of total fat per serving of the product.\
  Accessor: `per_serving`\
  Value Type: `number`
* **Daily Value**\
  DVs are the recommended amounts of total fat to consume or not to exceed each day\
  Accessor: `daily_value`\
  Value Type: `number`

Has a single value.

</details>

<details>

<summary>Saturated Fat</summary>

The amount of saturated fat in the product

Accessor: `saturated_fat`

**Subfields**

* **Per 100g**\
  The amount of saturated fat per 100g of the product\
  Accessor: `per_100g`\
  Value Type: `number`
* **Per Serving**\
  The amount of saturated fat per serving of the product\
  Accessor: `per_serving`\
  Value Type: `number`
* **Daily Value**\
  DVs are the recommended amounts of saturated fat to consume or not to exceed each day\
  Accessor: `daily_value`\
  Value Type: `number`

Has a single value.

</details>

<details>

<summary>Trans Fat</summary>

The amount of trans fat in the product

Accessor: `trans_fat`

**Subfields**

* **Per 100g**\
  The amount of trans fat per 100g of the product\
  Accessor: `per_100g`\
  Value Type: `number`
* **Per Serving**\
  The amount of trans fat per serving of the product\
  Accessor: `per_serving`\
  Value Type: `number`
* **Daily Value**\
  DVs are the recommended amounts of trans fat to consume or not to exceed each day\
  Accessor: `daily_value`\
  Value Type: `number`

Has a single value.

</details>

<details>

<summary>Cholesterol</summary>

The amount of cholesterol in the product

Accessor: `cholesterol`

**Subfields**

* **Per 100g**\
  The amount of cholesterol per 100g of the product\
  Accessor: `per_100g`\
  Value Type: `number`
* **Per Serving**\
  The amount of cholesterol per serving of the product\
  Accessor: `per_serving`\
  Value Type: `number`
* **Daily Value**\
  DVs are the recommended amounts of cholesterol to consume or not to exceed each day\
  Accessor: `daily_value`\
  Value Type: `number`

Has a single value.

</details>

<details>

<summary>Total Carbohydrate</summary>

The total amount of carbohydrates in the product

Accessor: `total_carbohydrate`

**Subfields**

* **Per 100g**\
  The amount of total carbohydrates per 100g of the product\
  Accessor: `per_100g`\
  Value Type: `number`
* **Per Serving**\
  The amount of total carbohydrates per serving of the product\
  Accessor: `per_serving`\
  Value Type: `number`
* **Daily Value**\
  DVs are the recommended amounts of total carbohydrates to consume or not to exceed each day\
  Accessor: `daily_value`\
  Value Type: `number`

Has a single value.

</details>

<details>

<summary>Dietary Fiber</summary>

The amount of dietary fiber in the product

Accessor: `dietary_fiber`

**Subfields**

* **Per 100g**\
  The amount of dietary fiber per 100g of the product\
  Accessor: `per_100g`\
  Value Type: `number`
* **Per Serving**\
  The amount of dietary fiber per serving of the product\
  Accessor: `per_serving`\
  Value Type: `number`
* **Daily Value**\
  DVs are the recommended amounts of dietary fiber to consume or not to exceed each day\
  Accessor: `daily_value`\
  Value Type: `number`

Has a single value.

</details>

<details>

<summary>Total Sugars</summary>

The total amount of sugars in the product

Accessor: `total_sugars`

**Subfields**

* **Per 100g**\
  The amount of total sugars per 100g of the product\
  Accessor: `per_100g`\
  Value Type: `number`
* **Per Serving**\
  The amount of total sugars per serving of the product\
  Accessor: `per_serving`\
  Value Type: `number`
* **Daily Value**\
  DVs are the recommended amounts of total sugars to consume or not to exceed each day\
  Accessor: `daily_value`\
  Value Type: `number`

Has a single value.

</details>

<details>

<summary>Added Sugars</summary>

The amount of added sugars in the product

Accessor: `added_sugars`

**Subfields**

* **Per 100g**\
  The amount of added sugars per 100g of the product\
  Accessor: `per_100g`\
  Value Type: `number`
* **Per Serving**\
  The amount of added sugars per serving of the product\
  Accessor: `per_serving`\
  Value Type: `number`
* **Daily Value**\
  DVs are the recommended amounts of added sugars to consume or not to exceed each day\
  Accessor: `daily_value`\
  Value Type: `number`

Has a single value.

</details>

<details>

<summary>Protein</summary>

The amount of protein in the product

Accessor: `protein`

**Subfields**

* **Per 100g**\
  The amount of protein per 100g of the product\
  Accessor: `per_100g`\
  Value Type: `number`
* **Per Serving**\
  The amount of protein per serving of the product\
  Accessor: `per_serving`\
  Value Type: `number`
* **Daily Value**\
  DVs are the recommended amounts of protein to consume or not to exceed each day\
  Accessor: `daily_value`\
  Value Type: `number`

Has a single value.

</details>

<details>

<summary>Sodium</summary>

The amount of sodium in the product

Accessor: `sodium`

**Subfields**

* **Per 100g**\
  The amount of sodium per 100g of the product\
  Accessor: `per_100g`\
  Value Type: `number`
* **Per Serving**\
  The amount of sodium per serving of the product\
  Accessor: `per_serving`\
  Value Type: `number`
* **Unit**\
  The unit of measurement for the amount of sodium\
  Accessor: `unit`\
  Value Type: `string`
* **Daily Value**\
  DVs are the recommended amounts of sodium to consume or not to exceed each day\
  Accessor: `daily_value`\
  Value Type: `number`

Has a single value.

</details>

<details>

<summary>Nutrients</summary>

The amount of nutrients in the product

Accessor: `nutrients`

**Subfields**

* **Name**\
  The name of nutrients of the product\
  Accessor: `name`\
  Value Type: `string`
* **Per 100g**\
  The amount of nutrients per 100g of the product\
  Accessor: `per_100g`\
  Value Type: `number`
* **Per Serving**\
  The amount of nutrients per serving of the product\
  Accessor: `per_serving`\
  Value Type: `number`
* **Unit**\
  The unit of measurement for the amount of nutrients\
  Accessor: `unit`\
  Value Type: `string`
* **Daily Value**\
  DVs are the recommended amounts of nutrients to consume or not to exceed each day\
  Accessor: `daily_value`\
  Value Type: `number`

Can have multiple values (is a list/array).

</details>


# Bank Account Details

Automatically parse bank account details and extract structured KYC data using the pre-trained Bank Account Details model template available in the Catalog.

Documentation for the data schema of the Bank Account Details model template.

## Bank Account Details Fields

Documentation for all fields present in the data schema.

Field *accessors* are used as the keys for accessing the values in the returned data. On the data schema interface, this is the "Field Name".

Field *value types* indicate how the value is returned by the API. On the data schema interface, this is the "Field Type".

<details>

<summary>Account Holder Name</summary>

The name of the account holder.

Accessor: `account_holder_name`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Account Number</summary>

The bank account number.

Accessor: `account_number`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Bank Name</summary>

The name of the bank.

Accessor: `bank_name`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Account Type</summary>

The type of bank account.

Accessor: `account_type`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>BIC</summary>

The Bank Identifier Code.

Accessor: `bic`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Branch Address</summary>

The address of the bank branch.

Accessor: `branch_address`

**Subfields**

* **Street**\
  The street address.\
  Accessor: `street`\
  Value Type: `string`
* **City**\
  The city.\
  Accessor: `city`\
  Value Type: `string`
* **State**\
  The state or province.\
  Accessor: `state`\
  Value Type: `string`
* **Postal Code**\
  The postal code.\
  Accessor: `postal_code`\
  Value Type: `string`
* **Country**\
  The country.\
  Accessor: `country`\
  Value Type: `string`

Has a single value.

</details>

<details>

<summary>IBAN</summary>

The International Bank Account Number.

Accessor: `iban`\
Value Type: `string`

Has a single value.

</details>


# Business Card

Automatically parse business cards and extract structured KYC data using the pre-trained Business Card model template available in the Catalog.

Documentation for the data schema of the Business Card model template.

## Business Card Fields

Documentation for all fields present in the data schema.

Field *accessors* are used as the keys for accessing the values in the returned data. On the data schema interface, this is the "Field Name".

Field *value types* indicate how the value is returned by the API. On the data schema interface, this is the "Field Type".

<details>

<summary>Name</summary>

The name of the person on the business card.

Accessor: `name`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Job Title</summary>

The job title of the person on the business card.

Accessor: `job_title`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Company</summary>

The company name on the business card.

Accessor: `company`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Phone Number</summary>

The phone number on the business card.

Accessor: `phone_number`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Email Address</summary>

The email address on the business card.

Accessor: `email_address`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Website</summary>

The website address on the business card.

Accessor: `website`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Address</summary>

The address on the business card.

Accessor: `address`\
Value Type: `string`

Has a single value.

</details>


# Payslip

Automatically parse payslips and extract structured salary data using the pre-trained Payslip model template available in the Catalog.

Documentation for the data schema of the Payslip model template.

## Payslip Fields

Documentation for all fields present in the data schema.

Field *accessors* are used as the keys for accessing the values in the returned data. On the data schema interface, this is the "Field Name".

Field *value types* indicate how the value is returned by the API. On the data schema interface, this is the "Field Type".

<details>

<summary>First Name</summary>

Employee's first name

Accessor: `first_name`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Last Name</summary>

Employee's last name

Accessor: `last_name`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Employee Address</summary>

Employee's address

Accessor: `employee_address`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Pay Period Start Date</summary>

Start date of the pay period

Accessor: `pay_period_start_date`\
Value Type: `date`

Has a single value.

</details>

<details>

<summary>Pay Period End Date</summary>

End date of the pay period

Accessor: `pay_period_end_date`\
Value Type: `date`

Has a single value.

</details>

<details>

<summary>Gross Pay</summary>

Gross pay amount

Accessor: `gross_pay`\
Value Type: `number`

Has a single value.

</details>

<details>

<summary>Net Pay</summary>

Net pay amount

Accessor: `net_pay`\
Value Type: `number`

Has a single value.

</details>

<details>

<summary>Deductions</summary>

List of deductions

Accessor: `deductions`

**Subfields**

* **Name**\
  Name of the deduction\
  Accessor: `name`\
  Value Type: `string`
* **Amount**\
  Amount of the deduction\
  Accessor: `amount`\
  Value Type: `number`

Can have multiple values (is a list/array).

</details>

<details>

<summary>Taxes</summary>

List of taxes

Accessor: `taxes`

**Subfields**

* **Name**\
  Name of the tax\
  Accessor: `name`\
  Value Type: `string`
* **Amount**\
  Amount of the tax\
  Accessor: `amount`\
  Value Type: `number`

Can have multiple values (is a list/array).

</details>

<details>

<summary>Employer Name</summary>

Name of the employer

Accessor: `employer_name`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Employer Address</summary>

Address of the employer

Accessor: `employer_address`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Social Security Number</summary>

Employee's Social Security Number

Accessor: `social_security_number`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Employee ID</summary>

Employee's ID

Accessor: `employee_id`\
Value Type: `string`

Has a single value.

</details>


# US Healthcare Card

Automatically parse US healthcare cards and extract structured healthcare provider data using the pre-trained US Healthcare Card model template available in the Catalog.

## US Healthcare Card

Documentation for the data schema of the US Healthcare Card model template.

## US Healthcare Card Fields

Documentation for all fields present in the data schema.

Field *accessors* are used as the keys for accessing the values in the returned data. On the data schema interface, this is the "Field Name".

Field *value types* indicate how the value is returned by the API. On the data schema interface, this is the "Field Type".

<details>

<summary>Company Name</summary>

The name of the company that provides the healthcare plan

Accessor: `company_name`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Plan Name</summary>

The name of the healthcare plan

Accessor: `plan_name`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Member Name</summary>

The name of the member covered by the healthcare plan

Accessor: `member_name`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Member ID</summary>

The unique identifier for the member in the healthcare system

Accessor: `member_id`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Issuer 80840</summary>

The organization that issued the healthcare plan

Accessor: `issuer_80840`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Dependents</summary>

The list of dependents covered by the healthcare plan

Accessor: `dependents`\
Value Type: `string`

Can have multiple values (is a list/array).

</details>

<details>

<summary>Group Number</summary>

The group number associated with the healthcare plan

Accessor: `group_number`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Payer ID</summary>

The unique identifier for the payer in the healthcare system

Accessor: `payer_id`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Rx BIN</summary>

The BIN number for prescription drug coverage

Accessor: `rx_bin`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Rx ID</summary>

The ID number for prescription drug coverage

Accessor: `rx_id`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Rx Grp</summary>

The group number for prescription drug coverage

Accessor: `rx_grp`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Rx PCN</summary>

The PCN number for prescription drug coverage

Accessor: `rx_pcn`\
Value Type: `string`

Has a single value.

</details>

<details>

<summary>Copayments</summary>

Copayments for covered services.

Accessor: `copayments`

**Subfields**

* **Service Name**\
  Service name\
  Accessor: `service_name`\
  Possible Values: `primary_care`, `emergency_room`, `urgent_care`, `specialist`, `office_visit`, `prescription`
* **Service Fees**\
  Service price\
  Accessor: `service_fees`\
  Value Type: `number`

Can have multiple values (is a list/array).

</details>

<details>

<summary>Enrollment Date</summary>

The date when the member enrolled in the healthcare plan

Accessor: `enrollment_date`\
Value Type: `date`

Has a single value.

</details>


# Mindee V1 Overview

{% hint style="danger" %}
**It is no longer possible to create V1 accounts,** [**sign up for V2**](https://app.mindee.com/signup?utm_source=docs) **instead.**

Existing **paying** V1 users will keep **all access** to their accounts, organizations, and APIs.
{% endhint %}

## What Is Mindee V1?

Mindee is a **powerful OCR software and an API-first platform** that helps developers automate applications' workflows by standardizing the document processing layer through data recognition for key information using computer vision and machine learning.

Our mission is to provide real-time, human-level accuracy in data extraction from paper and digital documents for our customers.

Mindee V1 is actively maintained, however no new functionalities are planned.

## Mindee V1 Products In a Nutshell

### Off-the-Shelf APIs

Ready-to-use, precise and efficient data handling. Our APIs offer seamless integration and high accuracy to streamline your workflows.

**Equivalent in Mindee V2:** create a model from a template.\
Unlike with V1, you can change the fields of the resulting model to better fit your needs.

### DocTI

docTI is an LLM-based solution that enables you to create within a few minutes an API to extract the precise data and fields you need from any form of document.

**Equivalent in Mindee V2:** create a model from scratch using our AI Assistant, or by simply uploading a sample file.\
Overall accuracy is improved compared to V1, with v2 models being more advanced. In addition, better tooling is provided, allowing finer tuning of parameters.


