Model Library
Browse and deploy state-of-the-art AI models through the DEVUP Gateway.
Browse and deploy state-of-the-art AI models through the DEVUP Gateway.
Browse and deploy state-of-the-art AI models through the DEVUP Gateway.

DeepSeek V4 Pro 0813 is the official release of V4 Pro, and the point at which the series flagship became a production agent rather than a preview. It keeps the architecture that defines the series — 1.6 trillion total parameters with 49 billion active per token, a one-million-token context window, and a hybrid attention stack built to make that window practical — and rebuilds the agentic behaviour on top of it. Long-horizon software engineering, terminal automation, security workflows, and tool orchestration all improve substantially over the preview, in several cases by a factor of several. Reasoning effort is a per-request setting with three levels, and the checkpoint ships with a speculative decoding module for faster generation. Released under the MIT license.

Gemini 3.7 Flash is Google's most capable Flash-tier model, built for coding and agentic work rather than for chat alone. It is natively multimodal in the widest sense available: text, images, video, audio, and PDFs all go in directly, against a one-million-token context window, with no separate extraction step in front of it. Thinking depth is a request-level setting with three levels, letting the same model serve a latency-critical pipeline and a long multi-step agent run. Google reports large gains over the previous Flash generation on issue resolution, production-ready code generation, complex document processing, and real-world business automation. On DEVUP AI it is the natural choice when your input is not plain text and your workflow has more than one step.

DeepSeek V4 Flash 0731 is the official release of V4 Flash, and it is the checkpoint where the series became genuinely agentic. Built on the same efficiency-first Mixture-of-Experts design — 284B total parameters with only 13B active per token and a one-million-token context window — it adds a substantially rebuilt agentic capability that lifts long-horizon coding, terminal automation, and tool-use scores far above the preview, in several cases past the much larger V4 Pro preview despite activating a fraction of its parameters. Reasoning effort is a per-request setting with three levels, and the checkpoint ships with a speculative decoding module for faster generation. Released under the MIT license, it is the model to reach for on DEVUP AI when an agent has to finish a long job, not just start one.

Step 3.7 Flash is an open-source multimodal reasoning model by StepFun with 198B total parameters (11B active) using Mixture of Experts. It accepts text and image inputs and features a 256K context window, selectable reasoning effort, tool calling, and agentic capabilities for coding and search workflows, scoring 80.9% on GPQA Diamond and 56.3% on SWE-bench Pro.

GLM-5.2 is Z-AI's latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a **solid 1M-token context**.

Kimi K2.7 Code is a coding-focused agentic model built upon Kimi K2.6. With substantial improvements on real-world long-horizon coding tasks, it strengthens end-to-end task completion across complex software engineering workflows while improving token efficiency, reducing thinking-token usage by approximately 30% compared with Kimi K2.6.

Qwen3.5-397B-A17B is Alibaba's most capable Qwen3.5 model, a Mixture-of-Experts architecture with 397B total parameters and 17B activated per token. It features a 262K token context window (extensible to 1M with YaRN), thinking/reasoning mode, tool calling with MCP integration, and support for 201 languages. Sets state-of-the-art results on reasoning, coding, math, and multimodal benchmarks.

The latest flagship reasoning model in the Qwen3 family. Further enhanced by multiple innovations like adaptive tool-use and advanced test-time scaling techniques

Qwen3.6-35B-A3B is Alibaba's latest flagship Mixture-of-Experts model, with 35B total parameters and only 3B activated per token (256 experts, 8 routed + 1 shared). Built on direct feedback from the community, Qwen3.6 prioritizes stability and real-world utility, offering developers a more intuitive, responsive, and genuinely productive coding experience.

NVIDIA Nemotron 3 Super is a hybrid Mixture-of-Experts (MoE) model engineered for highest compute efficiency and accuracy in multi-agent applications and specialized agentic systems. It is optimized to run many collaborating agents per application on a single GPU, delivering high accuracy for reasoning, tool use, and instruction following.

GLM-5 is an advanced, open-source large language model designed for developers tackling the toughest challenges. It excels at long-context reasoning, multi-step tool orchestration, and complex systems engineering, making it the ideal choice for powering sophisticated agents and applications that require high-level cognitive tasks.

GLM-4.7-Flash is a 30B-A3B MoE model. As the strongest model in the 30B class, GLM-4.7-Flash offers a new option for lightweight deployment that balances performance and efficiency.