Model Library
Browse and deploy state-of-the-art AI models through the DevUp Gateway.
Browse and deploy state-of-the-art AI models through the DevUp Gateway.
Browse and deploy state-of-the-art AI models through the DevUp Gateway.

DeepSeek V4 Pro 0813 is the official release of V4 Pro, and the point at which the series flagship became a production agent rather than a preview. It keeps the architecture that defines the series — 1.6 trillion total parameters with 49 billion active per token, a one-million-token context window, and a hybrid attention stack built to make that window practical — and rebuilds the agentic behaviour on top of it. Long-horizon software engineering, terminal automation, security workflows, and tool orchestration all improve substantially over the preview, in several cases by a factor of several. Reasoning effort is a per-request setting with three levels, and the checkpoint ships with a speculative decoding module for faster generation. Released under the MIT license.

Gemini 3.7 Flash is Google's most capable Flash-tier model, built for coding and agentic work rather than for chat alone. It is natively multimodal in the widest sense available: text, images, video, audio, and PDFs all go in directly, against a one-million-token context window, with no separate extraction step in front of it. Thinking depth is a request-level setting with three levels, letting the same model serve a latency-critical pipeline and a long multi-step agent run. Google reports large gains over the previous Flash generation on issue resolution, production-ready code generation, complex document processing, and real-world business automation. On DEVUP AI it is the natural choice when your input is not plain text and your workflow has more than one step.

DeepSeek V4 Flash 0731 is the official release of V4 Flash, and it is the checkpoint where the series became genuinely agentic. Built on the same efficiency-first Mixture-of-Experts design — 284B total parameters with only 13B active per token and a one-million-token context window — it adds a substantially rebuilt agentic capability that lifts long-horizon coding, terminal automation, and tool-use scores far above the preview, in several cases past the much larger V4 Pro preview despite activating a fraction of its parameters. Reasoning effort is a per-request setting with three levels, and the checkpoint ships with a speculative decoding module for faster generation. Released under the MIT license, it is the model to reach for on DEVUP AI when an agent has to finish a long job, not just start one.

Step 3.7 Flash is an open-source multimodal reasoning model by StepFun with 198B total parameters (11B active) using Mixture of Experts. It accepts text and image inputs and features a 256K context window, selectable reasoning effort, tool calling, and agentic capabilities for coding and search workflows, scoring 80.9% on GPQA Diamond and 56.3% on SWE-bench Pro.

GLM-5.2 is Z-AI's latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a **solid 1M-token context**.

Kimi K2.7 Code is a coding-focused agentic model built upon Kimi K2.6. With substantial improvements on real-world long-horizon coding tasks, it strengthens end-to-end task completion across complex software engineering workflows while improving token efficiency, reducing thinking-token usage by approximately 30% compared with Kimi K2.6.

● Qwen3-TTS-VoiceDesign is a voice design variant of Qwen3-TTS by Alibaba's Qwen team. Instead of selecting from preset voices, you describe the voice you want in natural language — and the model generates speech in that voice. Key capabilities: - Natural language voice control — describe any voice with free text (e.g. "a deep male voice with a calm, authoritative presence", "a young cheerful female with a warm and friendly tone") - 10 languages — English, Chinese, Japanese, Korean, German, French, Russian, Spanish, Italian, Portuguese - Streaming support — real-time PCM streaming - Multiple output formats — WAV, MP3, FLAC, PCM Built on the same 1.7B parameter architecture as Qwen3-TTS, using discrete multi-codebook language modeling and a custom 12Hz acoustic tokenizer for high-quality end-to-end speech synthesis.

Qwen3.5-397B-A17B is Alibaba's most capable Qwen3.5 model, a Mixture-of-Experts architecture with 397B total parameters and 17B activated per token. It features a 262K token context window (extensible to 1M with YaRN), thinking/reasoning mode, tool calling with MCP integration, and support for 201 languages. Sets state-of-the-art results on reasoning, coding, math, and multimodal benchmarks.

The latest flagship reasoning model in the Qwen3 family. Further enhanced by multiple innovations like adaptive tool-use and advanced test-time scaling techniques

Qwen3-TTS is an advanced text-to-speech model by Alibaba's Qwen team, delivering stable, expressive, and low-latency speech generation across 10 languages. Key capabilities: - 9 preset voices — Vivian, Serena, Uncle_Fu, Dylan, Eric, Ryan, Aiden, Ono_Anna, Sohee — covering diverse genders, ages, and accents - Voice cloning — clone any voice from a short (~3s) audio sample via the voice_id parameter - Instruction control — adjust tone, emotion, and speaking style with natural language (e.g. "speak slowly and calmly", "excited tone") - 10 languages — English, Chinese, Japanese, Korean, German, French, Russian, Spanish, Italian, Portuguese - Streaming support — real-time PCM streaming with ~97ms first-byte latency - Multiple output formats — WAV, MP3, FLAC, PCM Built on a 1.7B parameter architecture using discrete multi-codebook language modeling for end-to-end speech synthesis without cascading errors. Uses a custom 12Hz acoustic tokenizer that preserves paralinguistic information and environmental audio details.

Qwen3.6-35B-A3B is Alibaba's latest flagship Mixture-of-Experts model, with 35B total parameters and only 3B activated per token (256 experts, 8 routed + 1 shared). Built on direct feedback from the community, Qwen3.6 prioritizes stability and real-world utility, offering developers a more intuitive, responsive, and genuinely productive coding experience.

NVIDIA Nemotron 3 Super is a hybrid Mixture-of-Experts (MoE) model engineered for highest compute efficiency and accuracy in multi-agent applications and specialized agentic systems. It is optimized to run many collaborating agents per application on a single GPU, delivering high accuracy for reasoning, tool use, and instruction following.