Modelsgooglegemini-3.7-flash
providergoogle /

gemini-3.7-flash

263 DZD in 1313 DZD out/ 1M tokens

Gemini 3.7 Flash is Google's most capable Flash-tier model, built for coding and agentic work rather than for chat alone. It is natively multimodal in the widest sense available: text, images, video, audio, and PDFs all go in directly, against a one-million-token context window, with no separate extraction step in front of it. Thinking depth is a request-level setting with three levels, letting the same model serve a latency-critical pipeline and a long multi-step agent run. Google reports large gains over the previous Flash generation on issue resolution, production-ready code generation, complex document processing, and real-world business automation. On DEVUP AI it is the natural choice when your input is not plain text and your workflow has more than one step.

PublicJSON
gemini-3.7-flash
Capabilities
ToolsVisionReasoningStructured outputAudio
ArchitectureMoE
Context Window1M

Gemini 3.7 Flash

Overview

Gemini 3.7 Flash is Google's most capable Flash-tier model and the current workhorse of the Gemini 3 family. It sits between the deep-reasoning tier and the high-throughput lightweight tier, and it was tuned for a specific shape of work: multi-step agentic execution and production code, not single-turn question answering.

Two things define it in practice. It is natively multimodal across five input types — text, image, video, audio, and PDF — against a one-million-token context window, which removes the extraction and preprocessing layer most document and media pipelines are built around. And thinking depth is a request-level control with three levels, so latency-critical paths and long agent runs can share one model.

It is a direct iteration on the previous Flash generation, released weeks after it, with algorithmic improvements to the reasoning foundation rather than a new architecture.


⚠️ Sampling Parameters Are Not Supported

This model rejects temperature, top_p, and top_k. They were removed from the Gemini Flash line and are no longer accepted. This is the single most common cause of integration failure when migrating a working request from another model.

Most OpenAI-compatible clients send temperature by default. Remove it explicitly.

Prefilled model turns are also not supported. A messages array may not end with an assistant turn intended to be continued.

Output variance is controlled through the thinking level and through prompt instructions, not through sampling.


Thinking Levels

LevelBehaviourUse it for
lowReduced time-to-answerLatency-critical paths: real-time chat, incident pipelines, draft generation, fast data analysis
medium (default)BalancedGeneral work — the sensible default for most applications
highMaximum deliberationHard reasoning, long agent runs, complex document analysis

minimal is not a valid value on this model. Sending it returns an API validation error rather than silently falling back. If you are migrating a request from a model where minimal was accepted, this will fail loudly — which is the good case.

The thinking level is set with the thinking_level field. DEVUP AI forwards the complete request body upstream without stripping unknown fields, so it can be passed directly in your payload.


Multimodal Input

The widest input surface of any model in this catalogue.

Input typeSupported
TextYes
ImageYes
VideoYes
AudioYes
PDFYes
OutputText only

PDFs are read natively. There is no OCR step, no page splitting, and no text extraction library in front of the model — the document goes in as a document, and layout, tables, and figures are part of what the model sees. Combined with the context window, this means a full report or contract set can be reasoned over in one request.


Capabilities

CapabilityValue
input_typestext, image, video, audio, pdf
output_typestext
context_window1048576
max_output_tokens65536
reasoningNative — low, medium, high
streamingSupported
tool_callingSupported
structured_outputSupported — JSON schema
sampling_parametersNot supported — no temperature, top_p, top_k
requires_promptYes — text prompt required, media optional

Recommended Use Cases

  • Document intelligence — invoices, contracts, reports, forms. Native PDF input at full context is the capability that distinguishes this model most sharply.
  • Agentic coding and issue resolution — the workload it was built for, with large reported gains over the previous generation on debugging and production-ready code output.
  • Business process automation — multi-step real-world workflows, an area Google reports as substantially improved.
  • Web and front-end generation — more functional layouts and more complete applications in fewer prompts.
  • Audio and video understanding — transcription-adjacent work, meeting analysis, and video question answering without a separate speech model in the pipeline.
  • Knowledge-dense professional domains — finance, law, and biosciences, where the gain is in reading dense source material accurately rather than in recall.
  • Long-context retrieval — very high reported accuracy on long-context recall benchmarks.

Using Gemini 3.7 Flash on DEVUP AI

Base URL: https://api.devupai.com/v1 · Model ID: gemini-3.7-flash

Quick start — cURL

Note the absence of temperature. That is deliberate.

BASH
curl https://api.devupai.com/v1/chat/completions \
  -H "Authorization: Bearer $DEVUP_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemini-3.7-flash",
    "messages": [
      {
        "role": "user",
        "content": "Two services write to the same row without a transaction. Walk through the failure modes in order of likelihood and propose the smallest fix for each."
      }
    ],
    "thinking_level": "high",
    "max_tokens": 32768
  }'

PDF input — Python

The capability worth building around.

PYTHON
import os
import base64
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["DEVUP_API_KEY"],
    base_url="https://api.devupai.com/v1",
)

with open("supplier_contract.pdf", "rb") as handle:
    encoded = base64.b64encode(handle.read()).decode("utf-8")

response = client.chat.completions.create(
    model="gemini-3.7-flash",
    messages=[
        {
            "role": "user",
            "content": [
                {
                    "type": "file",
                    "file": {
                        "filename": "supplier_contract.pdf",
                        "file_data": f"data:application/pdf;base64,{encoded}",
                    },
                },
                {
                    "type": "text",
                    "text": (
                        "List every payment obligation in this contract as JSON, with fields: "
                        "clause_reference, party_owing, amount, currency, due_condition. "
                        "Quote the clause reference exactly as printed. "
                        "Use null for any field the document does not state — do not infer."
                    ),
                },
            ],
        }
    ],
    max_tokens=32768,
    extra_body={"thinking_level": "high"},
)

print(response.choices[0].message.content)

Instructing the model to return null rather than infer is not optional in document work. An invented figure is worse than a missing one, because nothing downstream can detect it.

Image input — Python

PYTHON
with open("invoice.png", "rb") as handle:
    encoded = base64.b64encode(handle.read()).decode("utf-8")

response = client.chat.completions.create(
    model="gemini-3.7-flash",
    messages=[
        {
            "role": "user",
            "content": [
                {
                    "type": "image_url",
                    "image_url": {"url": f"data:image/png;base64,{encoded}"},
                },
                {
                    "type": "text",
                    "text": (
                        "Extract every line item as JSON with fields: description, quantity, "
                        "unit_price, line_total. Report the currency exactly as printed."
                    ),
                },
            ],
        }
    ],
    max_tokens=16384,
)

print(response.choices[0].message.content)

Structured output with a schema

Rather than asking for JSON in prose and parsing defensively, constrain the response shape.

PYTHON
response = client.chat.completions.create(
    model="gemini-3.7-flash",
    messages=[{"role": "user", "content": ticket_text}],
    response_format={
        "type": "json_schema",
        "json_schema": {
            "name": "support_ticket",
            "strict": True,
            "schema": {
                "type": "object",
                "properties": {
                    "order_id": {"type": ["string", "null"]},
                    "issue_type": {
                        "type": "string",
                        "enum": ["missing_items", "damaged", "late", "other"],
                    },
                    "missing_items": {"type": "array", "items": {"type": "string"}},
                    "contact": {"type": ["string", "null"]},
                },
                "required": ["order_id", "issue_type", "missing_items", "contact"],
                "additionalProperties": False,
            },
        },
    },
    max_tokens=4096,
    extra_body={"thinking_level": "low"},
)

A schema-constrained response removes a whole class of parsing failure. Pair it with low thinking on high-volume extraction paths.

Node.js — DEVUP AI SDK

BASH
npm install devupai
JAVASCRIPT
import DevupAI from "devupai";

const client = new DevupAI({
  apiKey: process.env.DEVUP_API_KEY,
});

// No temperature, no top_p — this model does not accept them.
const response = await client.chat.completions.create({
  model: "gemini-3.7-flash",
  messages: [
    {
      role: "system",
      content:
        "You are a staff engineer reviewing a migration. Report only defects that would " +
        "cause data loss or downtime, each with a severity and the smallest safe fix.",
    },
    { role: "user", content: migrationPlan },
  ],
  max_tokens: 32768,
  thinking_level: "high",
});

console.log(response.choices[0].message.content);

Streaming with usage

PYTHON
stream = client.chat.completions.create(
    model="gemini-3.7-flash",
    messages=[{"role": "user", "content": "Design a retry policy for a webhook delivery system."}],
    max_tokens=32768,
    stream=True,
    stream_options={"include_usage": True},
)

for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
    if chunk.usage:
        print(f"\n\nTokens — in: {chunk.usage.prompt_tokens}, out: {chunk.usage.completion_tokens}")

Media inputs consume a substantial number of input tokens. On PDF, audio, and video workloads stream_options.include_usage is the only way to see how many before the request is over.

Delegating access with a scoped JWT

Document and media endpoints are frequently public-facing, which is exactly where an unrestricted key should never be. Issue a token limited to this model with an expiry and a spending limit:

BASH
curl -X POST "https://api.devupai.com/v1/scoped-jwt" \
  -H "Authorization: Bearer $DEVUP_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "api_key_name": "auto",
    "models": ["gemini-3.7-flash"],
    "expires_delta": 3600,
    "spending_limit": 200
  }'

The returned token is used exactly like an API key in the Authorization header. Requests for any other model, or past the expiry or spending limit, are rejected — a hard ceiling on what an uploaded file can cost you.


Benchmark Results

As reported by Google, with the previous Flash generation as the comparison point.

BenchmarkThis modelPrevious Flash generation
FrontierCode 1.1 Main43.6%34.4%
DeepSWE v1.165.3%49.0%
GDP.pdf (complex document processing)34.0%22.0%
AutomationBench30.4%17.0%
GDM-MRCR (long context)97.0%—
Harvey LAB-AA90.7%—
Code Arena, web development (Elo)1588—

The shape is consistent: the largest gains are in doing rather than knowing. Code issue resolution, document processing, and business workflow automation all improved by wide margins over a generation released only weeks earlier.


Best Practices

  • Remove temperature, top_p, and top_k from every request. This is the first thing to check when a working request from another model fails here.
  • Never end messages with an assistant turn. Prefilled model turns are not supported.
  • Match the thinking level to the path. low for extraction, classification, and latency-critical work; high for agent runs and dense documents; medium when unsure.
  • Never send minimal. It is rejected with a validation error.
  • Send documents as documents. Extracting text from a PDF before sending it discards the layout, tables, and figures the model would otherwise use.
  • Use schema-constrained structured output instead of asking for JSON in prose.
  • Instruct the model to return null rather than infer on any extraction task.
  • Watch input token counts on media requests. A batch of pages or a few minutes of audio consumes far more context than the surrounding text.

Limitations

  • Text output only. The model reads images, video, audio, and PDFs but generates only text. No image generation, no audio generation.
  • No sampling control. If your application depends on tuning temperature for output variance, that lever does not exist here.
  • Architecture is not disclosed. Google does not publish parameter counts, layer counts, or structural details for this model.
  • Some built-in tools are platform-specific. Capabilities tied to the model's native hosting environment are not part of the standard Chat Completions contract and should not be assumed available.
  • Standard foundation-model limitations apply, including hallucination. Validate any extracted figure before it reaches a system of record.
  • Not a safety layer. Apply your own moderation and validation before acting on model output in a production system.