Model Library
Browse and deploy state-of-the-art AI models through the DEVUP Gateway.
Browse and deploy state-of-the-art AI models through the DEVUP Gateway.
Browse and deploy state-of-the-art AI models through the DEVUP Gateway.

GLM-5.1 is Z-AI's next-generation flagship model for agentic engineering, with significantly stronger coding capabilities than its predecessor. It achieves state-of-the-art performance on SWE-Bench Pro and leads GLM-5 by a wide margin on NL2Repo (repo generation) and Terminal-Bench 2.0 (real-world terminal tasks).

MiMo-V2.5 is a native omnimodal model with strong agentic capabilities, supporting text, image, video, and audio understanding within a unified architecture. Built upon the MiMo-V2-Flash backbone and extended with dedicated vision and audio encoders, it delivers robust performance across multimodal perception, long-context reasoning, and agentic workflows.

MiMo-V2.5-Pro is an open-source Mixture-of-Experts (MoE) language model with 1.02T total parameters and 42B active parameters. It utilizes the hybrid attention architecture and 3-layers Multi-Token Prediction (MTP) introduced in [MiMo-V2-Flash](https://github.com/XiaomiMiMo/MiMo-V2-Flash).

Nemotron 3 Ultra is built for, frontier reasoning, orchestration, coding agents, deep research, and complex enterprise workflows. It delivers up to 5x faster inference and up to 30% lower cost for agentic workloads while supporting up to 1M token context.

Kimi K2.6 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration.

The latest flagship model in the Qwen family. State-of-the-art results across a comprehensive suite of benchmarks — including knowledge, reasoning, coding, instruction following, human preference alignment, agent tasks, and multilingual understanding.

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input and generating text output.

Efficient, MoE variant of Gemma 4. Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input and generating text output.

DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism that reduces training and inference cost while preserving quality in long-context scenarios. A scalable reinforcement learning post-training framework further improves reasoning, with reported performance in the GPT-5 class, and the model has demonstrated gold-medal results on the 2025 IMO and IOI. V3.2 also uses a large-scale agentic task synthesis pipeline to better integrate reasoning into tool-use settings, boosting compliance and generalization in interactive environments.

DeepSeek V4 Flash is an efficiency-first Mixture-of-Experts model built around a single idea: make a million tokens of context practical rather than theoretical. It holds 284B total parameters but activates only 13B per token, and pairs that with a hybrid attention stack designed specifically to cut the compute and memory cost of very long inputs. Three reasoning effort modes — non-think, high, and max — turn depth into a per-request setting, and the gap between them is dramatic: the same model answers instantly on routine work and competes with far larger systems on hard mathematics, competitive programming, and software engineering when given room to think. Released under the MIT license, it is the model to reach for on DEVUP AI when your input is long, your throughput matters, or both.

DeepSeek V4 Pro is the flagship of the DeepSeek V4 series: a Mixture-of-Experts model with 1.6 trillion total parameters, 49 billion of which activate per token, and a one-million-token context window. Its hybrid attention stack was designed specifically to make that window practical rather than nominal, cutting per-token compute to roughly a quarter and key-value cache to roughly a tenth of the previous generation at full context. Reasoning effort is a per-request control with three levels, and at its deepest setting the model reaches its strongest results on world knowledge, long-context retrieval, and multi-step agentic work. Released under the MIT license, it is the model to reach for on DEVUP AI when the input is enormous, the question is genuinely hard, or accuracy outweighs everything else.

Ling-3.0-flash-Fin is the first finance-enhanced model in Ant Group's Ling family, built by continuing the training of Ling-3.0-flash on high-quality financial data with input from financial institutions and domain experts. It carries 124 billion total parameters but activates only 5.1 billion per token — roughly four percent — which keeps inference light enough for the long agent runs financial research actually requires. Across a 256K context it connects retrieval, evidence review, calculation, modelling, and report preparation as one workflow rather than separate tasks, reconciling conflicting figures across annual reports, earnings releases, and filings. Released under the MIT licence, it retains the general reasoning, coding, and mathematics of the model it extends.