Instructions to use interfaze-ai/interfaze-1-lite with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use interfaze-ai/interfaze-1-lite with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("interfaze-ai/interfaze-1-lite", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Interfaze 1 Lite
Website ยท Docs ยท Run tasks ยท Blog ยท GitHub
Introduction
Interfaze 1 Lite is a mixture-of-architectures (MoA) model for developer workloads: reading documents, transcribing speech, locating objects and interface elements, and answering over all of it with structured output.
A single vision-language reasoning core works alongside a set of specialist architectures, each built for one kind of perception. The core reads the request, decides which specialists to run, and composes the answer from what they return. The whole model runs on one 80 GB GPU with no external services.
Key features
- Document understanding. Text, reading order, tables and layout from images, PDFs (up to 50 pages a call) and Word files, with a box and confidence for every line and word.
- Speech. Transcription with timestamps and speaker diarization. Long recordings are cut at pauses and decoded in batches: a 95-minute recording transcribes in about 90 seconds.
- Visual grounding. Open-vocabulary object detection with outlines, and GUI element grounding for computer-use agents.
- Structured output. Responses constrained to a JSON schema you supply, reading from any mix of text, images, documents and audio.
- Translation, forecasting and guardrails. Translation across 160+ languages, time-series forecasting from CSV or JSON, and safety checks on text and images.
- Multilingual reasoning. Science, math, SQL and general knowledge across 14+ languages, with a 131k-token context.
- Self-contained. One repository, one GPU, runs offline.
Model architecture
Interfaze 1 Lite is not one network. It is a reasoning core plus specialists, each chosen for the task it is best at, connected by tool calls.
| Component | Architecture | Role |
|---|---|---|
| Reasoning core | Hybrid-attention decoder with a vision encoder, FP8, 131k context | Plans, calls specialists, grounds boxes on a 0โ1000 grid, writes the answer |
| Document reader | Vision-language model trained for page reading | Text, reading order, tables and markdown |
| Line geometry | Text detector and recognizer | Every line's box and confidence |
| Layout | Document layout detector | Titles, paragraphs, tables and figures, with boxes |
| Speech | Encoder-decoder speech recognizer | Transcripts and timestamps in 99 languages |
| Diarization | Speaker segmentation and embedding pipeline | Who spoke when |
| Segmentation | Promptable segmentation model | Object outlines and masks |
| Forecasting | Time-series foundation model | Future values of a numeric series |
| Guardrails | Safety classifier | 14 text safety categories |
How the parts combine:
- OCR is two views of one page, stitched. The document reader supplies the text, and the line detector supplies the geometry. Each detected line takes the reader's words for it, so boxes are exact and text is complete.
- Speakers are attributed per word, by the largest overlap with each speaker's turns, then grouped into chunks.
- Detection and GUI grounding run on the reasoning core, which returns boxes on a 0โ1000 grid. Outlines come from the segmentation model.
- A run task skips planning. Naming one capability (
task="ocr","speech_to_text", โฆ) runs that specialist directly and returns its raw result.
Performance
| Benchmark | What it measures | Interfaze 1 Lite | Interfaze | GPT-5.4-Mini | Claude-Sonnet-4.6 | Gemini-3-Flash | Grok-4.3 |
|---|---|---|---|---|---|---|---|
| GPQA Diamond | Graduate-level science | 85.9 | 92.4 | 82.8 | 89.9 | 88.5 | 73.6 |
| MMMLU | Knowledge in 14 languages | 87.8 | 90.9 | 75.3 | 84.9 | 88.7 | 89.7 |
| MMMU-Pro | Multimodal reasoning | 73.2 | 71.1 | 40.4 | 46.3 | 67.6 | 68.7 |
| olmOCR-Bench | Document OCR | 83.8 | 85.7 | 80.1 | 73.9 | 75.3 | 81.9 |
| OCRBench v2 (English) | Text in images | 60.9 | 70.7 | 52.7 | 54.7 | 55.8 | 54.7 |
| RefCOCO (Acc@0.5) | Referring-expression grounding | 83.8 | 82.1 | โ | โ | โ | โ |
| VoxPopuli-Cleaned (WER โ) | Speech recognition | 3.01 | 2.4 | โ | โ | 4.0 | โ |
| SOB (value accuracy) | Structured output from text, images and audio | 81.5 | 80.5 | โ | 77.9 | 77.3* | โ |
| Spider 2.0-Lite (SQLite) | Text-to-SQL | 48.9 | 52.9 | 26.7 | 49.6 | 45.2 | 45.9 |
Interfaze 1 Lite was scored by us with each benchmark's official scorer. Every other score is from the Interfaze leaderboard. Higher is better except WER. *Gemini-3-Flash-Preview.
The benchmarks
- GPQA Diamond (all 198 questions). Graduate-level physics, chemistry and biology multiple choice, written to resist search. Lite scores 85.9, ahead of GPT-5.4-Mini and Grok-4.3, and strongest in physics.
- MMMLU (MMMLU-lite, all 19,950: 1,425 questions in each of 14 languages). MMLU translated by professional translators. Lite averages 87.8, ahead of Claude-Sonnet-4.6. The low-resource languages (Swahili, Yoruba, Bengali) are where it loses most.
- MMMU-Pro (all 1,730 questions per track, mean of standard and vision tracks). College-level questions that need the image, including a track where the question itself is inside the picture. Lite leads the board at 73.2.
- olmOCR-Bench (all 1,403 PDFs). Unit tests on real documents: arXiv math, old scans, tables, headers and footers, multi-column pages and long tiny text. Lite scores 83.8, with 91.9 on long tiny text and 88.8 on tables.
- OCRBench v2, English (all 7,400 English items). Recognition, referring, spotting, extraction, parsing, calculation, understanding and reasoning over text in images. Lite scores 60.9; its text spotting leads every general-purpose model on the board.
- RefCOCO (Acc@0.5). Find the one object a sentence describes ("the man in red on the left"). Lite's answer box scores 83.8, first on the board.
- VoxPopuli-Cleaned (all 628 clips). European Parliament speech, scored by word error rate after the benchmark's standard text normalisation. Lite's WER is 3.01%.
- SOB, the Structured Output Benchmark (all 5,324 records). Extract values into a JSON schema from text, images and audio; value accuracy counts exact field matches. Lite scores 81.5, second of 30 models, with 97% of responses valid JSON.
- Spider 2.0-Lite (the 135 SQLite tasks). Enterprise text-to-SQL over real schemas, scored by executing the query. Lite solves 48.9%, between the Claude models and Gemini-3-Flash.
Quickstart
To run Lite as an OpenAI-compatible server with Docker, see GitHub.
Requirements
- One 80 GB GPU with compute capability 8.9 or newer (Hopper, Ada). Tested on an H100.
ffmpegon the system, to decode audio.
hf download interfaze-ai/interfaze-1-lite requirements.txt --local-dir .
pip install -r requirements.txt
flash-linear-attention matters: without it, transformers runs the linear-attention layers as a plain PyTorch loop, and generation slows to minutes per page.
Using ๐ค Transformers
from transformers import AutoModel
model = AutoModel.from_pretrained("interfaze-ai/interfaze-1-lite", trust_remote_code=True)
answer = model.chat(
[{"role": "user", "content": "What is the total, and which item is highlighted?"}],
files=["receipt.jpg"],
)
print(answer["content"]) # the answer
print(answer["precontext"]) # each specialist's full result: [{"name", "result"}]
trust_remote_code=True is required: the model's routing and specialists are defined in this repository. Components load on first use, so a caller that only runs OCR never loads the others.
Each capability is also a method of its own:
| Method | Returns |
|---|---|
chat(messages, files=[...]) |
content (the answer) and precontext[] (each specialist's result) |
ocr(source, page_range=None, return_markdown=False) |
text, sections[] (one per page, with lines[].words[], four-corner bounds, average_confidence), width, height |
transcribe(audio, by_speaker=False, language="auto", word_timestamps=False) |
text and chunks[] with timestamps; each chunk carries a speaker with by_speaker |
detect(image, prompts, return_masks=False) |
detected_objects[] with label, bounds and polygon (and mask when asked) |
ground(image, prompts=None) |
gui_elements[] with type and bounds; with no prompts, every interactive element |
forecast(series, horizon) |
the next horizon points of a {date: value} series: timestamp[] and value[] |
moderate(text) |
output: "safe", or "unsafe" and the violated categories (S1โS14) |
Coordinates are pixels of the input: an image's own size, or a PDF page at 144 DPI.
Using the Interfaze API
The same model is served behind an OpenAI-compatible API. Get your API key from the Interfaze dashboard, then set model to interfaze-1-lite:
import { Interfaze } from "interfaze";
const interfaze = new Interfaze(); // reads INTERFAZE_API_KEY
const res = await interfaze.chat.completions.create({
model: "interfaze-1-lite",
messages: [{ role: "user", content: "Summarise the attached contract in three bullets." }],
});
Examples
Read a document, with boxes
doc = model.ocr("invoice.pdf", page_range=[1, 2])
print(doc["text"])
for page in doc["sections"]:
for line in page["lines"]:
box = line["bounds"]
print(page["page"], line["text"], box["top_left"], box["bottom_right"])
Transcribe a call and split it by speaker
call = model.transcribe("support_call.mp3", by_speaker=True)
for chunk in call["chunks"]:
start, end = chunk["timestamp"]
print(f"[{start:6.1f}โ{end:6.1f}] {chunk['speaker']}: {chunk['text']}")
Detect objects and outline them
found = model.detect("street.jpg", ["car", "bicycle", "traffic light"])
for obj in found["detected_objects"]:
print(obj["label"], obj["bounds"]["top_left"], obj["bounds"]["bottom_right"], len(obj.get("polygon", [])))
Ground interface elements for an agent
screen = model.ground("checkout.png", ["Add to cart button", "search box"])
for element in screen["gui_elements"]:
b = element["bounds"]
x = (b["top_left"]["x"] + b["bottom_right"]["x"]) / 2
y = (b["top_left"]["y"] + b["bottom_right"]["y"]) / 2
print(element["type"], "click at", (x, y))
Forecast a time series
weekly_sales = {
"2024-01-01": 412, "2024-01-08": 387, "2024-01-15": 524, "2024-01-22": 461,
"2024-01-29": 398, "2024-02-05": 542, "2024-02-12": 475, "2024-02-19": 401,
}
nxt = model.forecast(weekly_sales, horizon=4)
print(list(zip(nxt["timestamp"], nxt["value"])))
Check a message before it reaches your app
verdict = model.moderate("How do I make a weapon at home?")
print(verdict["output"]) # "safe", or "unsafe" and the violated codes on the next line
Extract structured data through the API
from openai import OpenAI
client = OpenAI(base_url="https://api.interfaze.ai/v1", api_key="sk_...")
res = client.chat.completions.create(
model="interfaze-1-lite",
messages=[{"role": "user", "content": [
{"type": "text", "text": "Extract the vendor, date and total."},
{"type": "image_url", "image_url": {"url": "https://example.com/receipt.jpg"}},
]}],
response_format={"type": "json_schema", "json_schema": {"name": "receipt", "schema": {
"type": "object",
"properties": {"vendor": {"type": "string"}, "date": {"type": "string"}, "total": {"type": "number"}},
"required": ["vendor", "date", "total"],
}}},
)
print(res.choices[0].message.content)
Limitations
- Generation through transformers is correct but slower than a serving engine with paged attention and batching. For throughput, use the Interfaze API or serve the model with a batching engine.
chatdoes not take a response schema; ask for JSON in the prompt, or use the API'sresponse_format.- A dense document page can take a minute or more to read on the transformers path.
- Memory: processing large PDFs can spike in significant use of CUDA memory.
Thank you
We are grateful for the inspiration from these models and the teams behind them: Qwen3.8 27B from the Qwen team, Chandra OCR 2 from Datalab, Whisper large-v3 turbo from OpenAI, speaker-diarization-community-1 from pyannote, SAM 2.1 and Llama Guard 3 from Meta, TimesFM 2.5 from Google Research, and PP-OCRv5 detection, PP-OCRv5 recognition and PP-DocLayout from PaddlePaddle. Thanks also to the open-source projects that run them: vLLM, Hugging Face Transformers, pyannote.audio, PaddleOCR, SAM 2 and TimesFM.
- Downloads last month
- 6
