Model Library
Browse and deploy state-of-the-art AI models through the DEVUP Gateway.
Browse and deploy state-of-the-art AI models through the DEVUP Gateway.
Browse and deploy state-of-the-art AI models through the DEVUP Gateway.

Gemini 2.5 Pro is Google's frontier-level multimodal model, engineered for highly complex reasoning, advanced coding, and massive-scale data analysis. Featuring an unprecedented 2M context window and hybrid thinking capabilities, it excels at solving the most demanding enterprise and scientific tasks.

Gemini 2.5 Flash is Google's latest thinking model, designed to tackle increasingly complex problems. It's capable of reasoning through their thoughts before responding, resulting in enhanced performance and improved accuracy. Gemini 2.5 Flash: best for balancing reasoning and speed.

The next generation of Anthropic's fastest and most cost-effective model, optimal for use cases where speed and affordability matter.

Claude Sonnet 4.6 delivers frontier intelligence at scale—built for coding, agents, and enterprise workflows.

Anthropic's most capable production model yet, advancing performance across coding, enterprise workflows, and long-running agentic tasks.

DeepSeek-V3-0324, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token, an improved iteration over DeepSeek-V3.

The DeepSeek R1 model has undergone a minor version upgrade, with the current version being DeepSeek-R1-0528.

DeepSeek-V3.1 is post-trained on the top of DeepSeek-V3.1-Base, which is built upon the original V3 base checkpoint through a two-phase long context extension approach, following the methodology outlined in the original DeepSeek-V3 report. We have expanded our dataset by collecting additional long documents and substantially extending both training phases. The 32K extension phase has been increased 10-fold to 630B tokens, while the 128K extension phase has been extended by 3.3x to 209B tokens. Additionally, DeepSeek-V3.1 is trained using the UE8M0 FP8 scale data format to ensure compatibility with microscaling data formats.

DeepSeek V3 is the model that established the architecture the rest of the line was built on. A Mixture-of-Experts design with 671B total parameters and 37B active per token, it pairs Multi-head Latent Attention with DeepSeekMoE, and introduced two ideas that have since become standard: an auxiliary-loss-free load balancing strategy, and a multi-token prediction training objective that doubles as a speculative decoding path at inference. It was the first model of this scale trained in FP8 mixed precision, on 14.8 trillion tokens, and its post-training distilled reasoning behaviour from a long chain-of-thought model without exposing a reasoning trace. It answers directly, holds a 128K context, and remains a capable general-purpose choice on DEVUP AI.