Model Library
Browse and deploy state-of-the-art AI models through the DEVUP Gateway.
Browse and deploy state-of-the-art AI models through the DEVUP Gateway.
Browse and deploy state-of-the-art AI models through the DEVUP Gateway.

Gemini 3.7 Flash is Google's most capable Flash-tier model, built for coding and agentic work rather than for chat alone. It is natively multimodal in the widest sense available: text, images, video, audio, and PDFs all go in directly, against a one-million-token context window, with no separate extraction step in front of it. Thinking depth is a request-level setting with three levels, letting the same model serve a latency-critical pipeline and a long multi-step agent run. Google reports large gains over the previous Flash generation on issue resolution, production-ready code generation, complex document processing, and real-world business automation. On DEVUP AI it is the natural choice when your input is not plain text and your workflow has more than one step.

Step 3.7 Flash is an open-source multimodal reasoning model by StepFun with 198B total parameters (11B active) using Mixture of Experts. It accepts text and image inputs and features a 256K context window, selectable reasoning effort, tool calling, and agentic capabilities for coding and search workflows, scoring 80.9% on GPQA Diamond and 56.3% on SWE-bench Pro.

Kimi K2.7 Code is a coding-focused agentic model built upon Kimi K2.6. With substantial improvements on real-world long-horizon coding tasks, it strengthens end-to-end task completion across complex software engineering workflows while improving token efficiency, reducing thinking-token usage by approximately 30% compared with Kimi K2.6.

Qwen3.5-397B-A17B is Alibaba's most capable Qwen3.5 model, a Mixture-of-Experts architecture with 397B total parameters and 17B activated per token. It features a 262K token context window (extensible to 1M with YaRN), thinking/reasoning mode, tool calling with MCP integration, and support for 201 languages. Sets state-of-the-art results on reasoning, coding, math, and multimodal benchmarks.

Qwen3.6-35B-A3B is Alibaba's latest flagship Mixture-of-Experts model, with 35B total parameters and only 3B activated per token (256 experts, 8 routed + 1 shared). Built on direct feedback from the community, Qwen3.6 prioritizes stability and real-world utility, offering developers a more intuitive, responsive, and genuinely productive coding experience.

MiMo-V2.5 is a native omnimodal model with strong agentic capabilities, supporting text, image, video, and audio understanding within a unified architecture. Built upon the MiMo-V2-Flash backbone and extended with dedicated vision and audio encoders, it delivers robust performance across multimodal perception, long-context reasoning, and agentic workflows.

Nemotron 3 Ultra is built for, frontier reasoning, orchestration, coding agents, deep research, and complex enterprise workflows. It delivers up to 5x faster inference and up to 30% lower cost for agentic workloads while supporting up to 1M token context.

Kimi K2.6 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration.

Efficient, MoE variant of Gemma 4. Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input and generating text output.

DeepSeek V4 Flash Vision Exp is an experimental vision-enabled build of V4 Flash, and it takes an unusual approach to image cost: every image is capped at 384 tokens regardless of its resolution. A two-megapixel scan and a twenty-five-megapixel one bill identically, because both are resized to the same working size before inference. That turns document processing from an open-ended cost into arithmetic you can do before sending the request. It accepts up to six hundred images per call, matches the text model it is built on across agents, reasoning, and world knowledge, and holds the same million-token context. On DEVUP AI it is the choice when the workload is pages rather than paragraphs.

Qwen3.5-122B-A10B is a large Mixture-of-Experts model from Alibaba's Qwen3.5 series with 122B total parameters and 10B activated per token. It features a 262K token context window (extensible to 1M with YaRN), thinking/reasoning mode, tool calling, and support for 201 languages. Excels at complex reasoning, coding, multimodal understanding, and agentic tasks with the efficiency of sparse activation.

Qwen3.8-27B is the compact member of Alibaba's most capable open model generation, and it makes an unusual argument: that a dense 27-billion-parameter model can compete with systems a hundred times its size on the work most applications actually do. It is a native vision-language model reading images and hour-scale video, with a hybrid attention stack that mixes gated linear attention with full attention across 64 layers, and a 262K native context extensible far beyond it. Thinking is controllable along three separate axes — on or off, three depth levels, and whether reasoning carries forward across turns. It leads its class on document intelligence and structured extraction, and on DEVUP AI it is the choice when the work is high-volume and the documents are real.