providerthinkingmachines /

Inkling-Small

203 DZD in 504 DZD out 40.6 DZD cached/ 1M tokens

Inkling-Small is a Mixture-of-Experts transformer with 276B total parameters, 12B active, trained on NVIDIA GB300 NVL72 systems. Like Inkling, it features native reasoning over audio and images, variable thinking effort

Publicfp8JSON
Inkling-Small
Capabilities
ToolsVisionReasoningStructured output
ArchitectureTransformer
Context Window524K

license: apache-2.0 license_link: https://www.apache.org/licenses/LICENSE-2.0 pipeline_tag: image-text-to-text tags:

  • conversational
  • image-text-to-text
  • audio-text-to-text
  • moe library_name: transformers

Inkling

1. General Information

Inkling-Small is a general-purpose multimodal model that accepts text, image, and audio inputs and generates text outputs. It is intended for use in English and other languages, and across multiple coding languages. The model is designed to be used by developers building AI-powered applications, including agentic and tool-use systems, coding assistants, chatbots, and retrieval-augmented generation systems, and is suitable for general-purpose conversational use, instruction-following, and other natural language and multimodal tasks. It is released with open weights to support research, fine-tuning, and integration into third-party products by downstream developers.

Languages: English, with general multilingual capabilities across other languages.


2. Model Properties

PropertyValue
Model typeMultimodal autoregressive transformer
Architecture typeA 42-layer decoder-only transformer with a sparse Mixture-of-Experts (MoE) feed-forward backbone: each token is routed to 6 of 256 experts, plus 2 shared experts active on every token. Attention is a hybrid of local and global layers. The model is natively multimodal — images are encoded via a hierarchical patch encoder, and audio via discrete token encoding — with all modalities projected into a shared hidden space and processed jointly by the decoder.
Parameters276B total, 12B active
Numerics supportBF16 and NVFP4

Input Modalities

Inkling-Small accepts text, image, and audio inputs:

  • Text: UTF-8 encoded text
  • Image: Any pixel-based image input. For optimal performance, each image dimension should be between 40px to 4096px.
  • Audio: WAV format, sampled at 16kHz. For optimal performance, audio length should ideally be under 2 minutes.

Output Modalities

Inkling-Small generates output as UTF-8 encoded text.


3. Evaluations

Category / BenchmarkInkling-SmallQwen3.5 397B-A17BMiMo V2.5Minimax M2.7DeepSeek V4 FlashNemotron 3 UltraInklingClaude 4.5 HaikuGemini 3.5 Flash-LiteGPT 5.6 Luna
Model Info
AA Index (v4.1)40.0%34.0%37.0%38.0%40.0%38.0%41.0%30.0%36.0%49.0%
Params (B) (activated / total)12 / 27617 / 39715 / 31010 / 23013 / 28455 / 55041 / 975–––
Agentic (coding)
SWEBench Verified80.2%76.4%71.0%79.9%79.0%70.7%77.6%73.3%75.0%93.0%
SWEBench Pro (public)55.9%50.9%56.1%56.2%52.6%46.4%54.3%39.5%54.2%62.7%
Terminal Bench 2.1 (best harness)64.7%51.3%63.7%55.4%61.8%56.4%63.8%44.2%54.0%82.5%
SciCode48.7%42.0%43.1%47.0%44.9%39.9%46.1%43.3%40.9%50.0%
Agentic (general)
GDPval-AA v212699621145115911891164123891111391530
MCP Atlas (public / all)79.6/79.2%74.2%/––49.4%/–69.0%/–47.4/44.7%78.8/76.0%41.2/40.2%79.8/76.8%77.0/75.0%
Tau 3 Banking15.5%13.4%6.6%8.9%22.9%13.8%23.7%9.1%16.5%24.3%
BrowseComp (with context management)77.4%78.6%–76.3%73.2%63.0%77.1%––84.0%
Toolathlon Verified54.4%40.7%49.1%47.5%50.9%34.3%45.5%26.9%57.1%67.9%
AA-Briefcase917–––833870839612––
Reasoning (general)
GPQA Diamond89.5%89.3%84.9%87.4%89.4%86.7%87.2%67.2%83.8%89.5%
HLE (text only)31.6%27.3%25.2%28.1%32.1%26.6%29.7%9.7%17.5%35.6%
HLE (with tools)47.8%48.3%40.0%40.3%45.1%37.4%46.0%17.8%42.5%48.9%
AIME 202695.5%93.3%93.6%87.7%95.8%94.2%97.1%85.1%82.2%97.6%
HMMT Feb 202690.2%87.9%82.6%71.2%93.9%78.8%86.3%66.7%63.6%98.5%
CritPt8.3%1.7%3.7%0.6%7.1%3.1%5.4%0.0%0.0%20.6%
Reasoning (abstract)
ARC-AGI-184.0%–––––79.5%47.7%–87.7%
ARC-AGI-240.1%–––––36.5%4.0%–47.6%
Factuality
SimpleQA Verified20.6%26.0%16.1%13.5%34.1%32.4%43.9%5.9%44.1%41.7%
AA Omniscience (index)-9.0-29.8-9.30.7-22.9-1.02.1-4.26.9-11.6
Chat
IFBench82.2%78.8%67.1%75.7%79.2%81.4%79.8%54.3%78.6%67.3%
Global-MMLU-Lite86.7%90.0%83.5%83.9%88.4%85.6%88.7%83.4%89.4%88.7%
Safety
StrongREJECT98.4%99.4%99.3%99.4%97.4%98.7%98.6%98.6%97.6%98.7%
FORTRESS (adversarial)71.6%77.3%64.8%86.3%32.0%77.6%78.0%91.3%70.7%83.8%
FORTRESS (benign)96.9%95.4%94.6%90.1%99.2%90.6%95.9%94.1%95.5%97.8%
Vision
MMMU Pro (Standard 10)74.0%77.3%75.4%–––73.5%58.6%79.0%78.6%
Charxiv RQ (original / with python)77.4/81.3%80.8%/–81.0%/––––78.1/82.0%57.4%/–70.0%/–81.4%/–
Audio
Audio MC54.9%–30.4%–––56.6%–33.6%–
MMAU77.0%–73.6%–––77.2%–75.2%–
VoiceBench90.1%–86.4%–––91.4%–85.9%–

Notes

  • Inkling-Small is evaluated against open- and closed-weights models across the full eval suite. Activated and total parameters are given for scale; a dash means the score was not available at the time of writing.
  • SWEBench Verified: Inkling and Inkling-Small's numbers are reported using a bash-only harness. Self-reported numbers are used for external models.
  • Terminal Bench 2.1: Inkling and Inkling-Small's numbers are reported using an internal coding harness. A small number of solutions were found to be contaminated from web search and were assigned a score of 0. Self-reported numbers are used for external models where available; otherwise, performance is reported using the internal harness.
  • Audio MC: Other models were evaluated internally since they are not on the official leaderboard.
  • VoiceBench: VoiceBench uses rule-based, hard-coded string matching for grading, making the evaluation sensitive to output-formatting differences. A system message instructing models to follow the expected answer format was therefore added.
  • HLE with tools: Minimax M2.7, Claude 4.5 Haiku, Gemini 3.5 Flash-Lite, and GPT 5.6 Luna were benchmarked using an internal harness.