Platform – Custom Vision Language Models (VLM)
To access the Ximilar VLM Platform, first register at Ximilar App to get your API token. This service is currently in beta and only available to selected users.
Vision Language Models (VLM) by Ximilar enable you to train custom vision-language models on your own images. Unlike simple image classification, a custom VLM can:
- generate structured outputs like JSON/YAML/XML/CSV responses with explanations
- analyze multiple images (or video frames) at once
- accept meta data to guide the analysis
- reason over several turns and call your own tools (agentic mode)
- produce embeddings for visual and multimodal search (retrieval mode)
Task Modes
Every VLM task and dataset has a mode. The mode decides what kind of data you annotate, which model is trained,
and what the trained model produces. A task can only be connected to datasets of the same mode, and the mode cannot be changed after creation.
| Mode | What you annotate | Trained model (default) | Prompts required | Training unit |
|---|---|---|---|---|
instruction | One or more images plus variable values that fill a result template | LiquidAI/LFM2.5-VL-450M (LoRA) | system prompt, user prompt, result template | samples |
agentic | A multi-turn conversation of steps: user messages, assistant thoughts, tool calls, tool results and a final answer | LiquidAI/LFM2.5-VL-450M (LoRA) | system prompt, user prompt | samples |
retrieval | Anchor (query) and document (candidate) items linked by relevance edges (positive / negative) | Qwen/Qwen3-VL-Embedding-2B (embedding model, LoRA) | none | anchors with at least one positive document |
Instruction mode
Instruction fine-tuning teaches the model to follow an instruction and answer in a fixed format.
Each dataset defines a result template with {{variable}} placeholders and a set of variables that describe the output schema.
Every sample holds images and the annotated variable values. During training the model learns to render the template from the visual content.
Agentic mode
Agentic fine-tuning teaches the model to solve a task over several turns. You define tools (functions with a JSON Schema of arguments) and annotate each sample as an ordered list of steps: the user question with images, the assistant's thoughts, tool calls, the tool results and the final answer.
Retrieval mode
Retrieval training produces an embedding model for search and RAG systems. Each sample is a small relevance graph:
one or more anchor items (the query side, e.g. a photo or a text query), one or more document items (the candidates),
and a relevance edge between every anchor and every document. Relevance 1.0000 marks a positive (matching) document, 0.0000 a negative one.
Key Concepts
- Prompt: Reusable named text block (
system,userortemplate) shared across tasks and datasets - Task: Defines the model to train. Has a
mode, references system/user prompts, connects to datasets - Model: A trained version of a task with stored weights, status and metrics
- Dataset: Collection of training samples with the same
modeas the task. Holds prompts, the result template and variables - Sample: One training example. Inherits its mode from the dataset. Samples can be flagged as
testfor evaluation - Variable (instruction): Schema definition of one output value (type, constraints, choices)
- Tool (agentic): A function the model may call, defined by a name, a description and a JSON Schema of parameters
- Step (agentic): One turn of a sample conversation (
role+type+content, optionally images) - Retrieval Item (retrieval): An
anchorordocumentmade of text and/or images - Retrieval Edge (retrieval): The relevance (
0.0000or1.0000) between one anchor and one document
Use Cases
- Product Description Generation: Generate structured product descriptions from images
- Quality Grading: Analyze items and generate quality grades with explanations
- Structured Data Extraction: Extract specific data points from images or documents in JSON format
- Visual Agents: Let the model look up a barcode, crop an object or query a database before answering
- Multimodal Search: Train an embedding model that matches product photos to catalogue entries or text queries
Workflow
The steps are the same for every mode; only the sample content differs.
- Create the prompts you need with Create Prompt (none for retrieval).
- Create a Task with the chosen
mode. - Create a Dataset with the same
modeand connect it to the task. - Upload your images with the Upload Training Image endpoint of the recognition API. VLM samples reference images by their
id. - Create samples and fill them:
- instruction: add images, create variables and set variable values
- agentic: create tools, then create steps and attach images to the user step
- retrieval: create anchor and document items, attach images, then create relevance edges
- Check your data with the validate endpoints, then train the task.
- Once a model is
TRAINED, run inference with async requests.
Training requirements. A task needs at least 20 samples across its datasets (for retrieval tasks: at least 20 anchors that have a positive document).
This is only the technical minimum: to fine-tune a VLM that performs well you will typically need a few thousand samples, so treat 20 as a smoke test, not a production dataset.
A sample, step or retrieval item can hold at most 10 images and 10 detection objects.
All VLM endpoints accept an optional ?workspace=WORKSPACE_ID query parameter; without it your default workspace is used.
API Reference
The endpoint reference is split by resource:
- Prompts: reusable system prompts, user prompts and result templates
- Tasks: create tasks, connect datasets, validate and train
- Models: trained model versions, their status, metrics and downloads
- Requests: run inference with async requests
- Datasets: create, validate and import training datasets
- Samples: create and manage training samples (shared by all modes)
- Instruction Datasets: variables, sample images, detection objects and variable values
- Agentic Datasets: tools and multi-turn steps
- Retrieval Datasets: anchor and document items and relevance edges
All Endpoints
https://api.ximilar.com/vlm/v2/prompt/
https://api.ximilar.com/vlm/v2/prompt/__PROMPT_ID__/
Video and Audio Attachments
Besides images, samples (instruction), steps (agentic, on user and tool_result steps only) and retrieval items can hold media assets
(video or audio files uploaded through the Media API). The endpoints share one shape:
| Parent | Endpoints |
|---|---|
| sample | POST /v2/sample/{id}/add-media/, POST .../remove-media/, GET .../sample-media/, GET, PATCH .../sample-media/{media_id}/, POST .../reorder-media/ |
| step | POST /v2/step/{id}/add-media/, POST .../remove-media/, GET .../step-media/, GET, PATCH .../step-media/{media_id}/, POST .../reorder-media/ |
| retrieval item | POST /v2/retrieval_item/{id}/add-media/, POST .../remove-media/, GET .../item-media/, GET, PATCH .../item-media/{media_id}/, POST .../reorder-media/ |
add-mediabody:{"items": [{"asset_id": "__ASSET_ID__", "text": "optional label", "start_ms": 0, "end_ms": 5000}]};start_msandend_msselect a clip and must be given togetherremove-mediabody:{"attachment_ids": ["__ATTACHMENT_ID__"]}reorder-mediabody: the full new order of all attachments of the parent,{"items": [{"attachment_type": "image" | "detection_object" | "media", "attachment_id": "..."}]}
Media attachments can be stored, annotated and previewed, but the current models cannot train on them yet.
A task whose datasets contain video or audio attachments is rejected by Train Task with error_type: unsupported_media.
Using Different Workspace
When making an API request, the default workspace associated with the user's API token is used. To access data or upload to a different workspace, specify the workspace in the URL or JSON payload.
# Get all samples from a specific workspace
https://api.ximilar.com/vlm/v2/sample/?workspace=WORKSPACE_ID
# Create a sample in a specific workspace
curl -XPOST \
-H 'Authorization: Token __API_TOKEN__' \
-H 'Content-Type: application/json' \
-d '{
"dataset": "__DATASET_ID__",
"workspace": "WORKSPACE_ID"
}' \
https://api.ximilar.com/vlm/v2/sample/
from ximilar.client.vlm import VLMClient
# Initialize client with specific workspace
client = VLMClient(
token="__API_TOKEN__",
workspace="WORKSPACE_ID"
)