Modelsopenaigpt-6.1-sol
provideropenai /

gpt-6.1-sol

700 DZD in 3500 DZD out 35 DZD cached/ 1M tokens

GPT-6.1 Sol adds two things to the tier below OpenAI's flagship, and takes one away. It reads images, which the model it replaces did not. Its knowledge runs ten days later. And it no longer accepts the lowest reasoning settings — none and minimal are rejected, which makes deliberation mandatory and breaks any integration written against its predecessor's full range. A 1,050,000-token window with 922,000 available for input and 128,000 for output, five effort levels from low to max, and OpenAI's own documentation telling you to compare it against the flagship on your own tasks rather than take the positioning on trust.

PublicJSONStreaming
gpt-6.1-sol
Capabilities
ToolsVisionReasoningStructured output
ArchitectureProprietary
Context Window1M

GPT-6.1 Sol

Near-flagship capability a tier down — with one setting removed that will break your migration.

Released 29 September 2026 — model page


none and minimal Are Rejected

The first thing to check before pointing existing code at this model, stated plainly in OpenAI's own documentation:

reasoning.effort supports low, medium (default), high, xhigh, and max. The none and minimal reasoning efforts are not supported.

The model it replaces accepted none. This one does not.

Which means reasoning is mandatory here. There is no configuration that skips deliberation entirely — low is the floor, and low still thinks.

Three consequences.

Existing requests will fail. Code that sets reasoning.effort: "none" — a reasonable setting for classification, extraction, or formatting on the previous model — is rejected rather than silently upgraded.

Your fast path needs rebuilding. Whatever you were routing to none now goes to low, or to a different model.

And max_tokens must cover reasoning plus answer, always. On the previous model, none gave you a budget that was entirely answer. That configuration no longer exists.

Audit your request builders before switching. This is a one-line change to find and a silent production failure if you miss it.


Vision Is New Here

Input: text and images. Output: text.

The model this replaces was text-only. Adding image input inside a point release is unusual — it is the kind of capability that normally waits for a whole-number generation.

What it enables at this tier. Document extraction, screenshot analysis, chart reading, and computer use — all without reaching for the flagship above it.

And computer use is named directly in OpenAI's positioning: complex coding, computer use, and professional work. A model driving an interface has to read the screen before it acts, which is vision and decision in one capability.


922,000 In, 128,000 Out

Three figures, and the middle one is a separate cap rather than a remainder.

Context window1,050,000 tokens
Maximum input922,000 tokens
Maximum output128,000 tokens

922,000 is a hard ceiling on input, not what remains after output. You cannot send a million tokens and ask for a short answer.

128,000 is a hard ceiling on output, not a share you enlarge by sending less.

They sum to the headline figure, which makes the total easy to satisfy and the individual caps easy to violate. A request of 950,000 input and 50,000 output comes in under 1,050,000 and still fails.

Check each separately.


OpenAI Tell You to Compare

A sentence in their own documentation worth reading twice:

Compare it with Astra on your tasks to assess the tradeoff between quality and cost.

A vendor instructing you to test rather than trust the positioning.

Which is the right advice and an unusual place to find it. "Near-flagship performance at a lower cost" is a claim about an average across tasks. Whether it holds on yours is a different question, and the only way to answer it is to run both.

The comparison is cheap. Same interface, same request shape, one changed model identifier. Twenty of your real tasks through each will tell you more than any published figure.


Specifications

Model IDopenai/gpt-6.1-sol
Context window1,050,000 tokens
Maximum input922,000 tokens
Maximum output128,000 tokens
Reasoninglow, medium (default), high, xhigh, max
Not supportednone, minimal
Reasoning qualityHighest
SpeedFast
InputText, image
OutputText
Knowledge cutoff30 April 2026
Fine-tuningListed as supported
WeightsClosed
Released29 September 2026
DeveloperOpenAI

A single snapshot. gpt-6.1-sol is a dotted version rather than a dated checkpoint — there is no gpt-6.1-sol-2026-09-29 alongside it.

A priority-tier variant exists for paths where latency outweighs everything else.

Tools through the Responses API: web search, image generation, code interpreter, and a hosted shell.


Capabilities

CapabilityValue
input_typestext, image
output_typestext
context_window1048576
max_input_tokens~922,000
max_output_tokens~128,000
reasoningMandatory — low through max
reasoning_disableNot supported
streamingSupported
tool_callingSupported
structured_outputSupported
fine_tuningListed
requires_promptYes — text prompt required, image optional

Ten Days of Knowledge

A small detail with a real consequence.

ModelKnowledge cutoff
The model it replaces20 April 2026
GPT-6.1 Sol30 April 2026

Ten days further forward.

Which sounds trivial and is not, on a fast-moving topic. A library release, a regulatory change, or a product launch inside that window is something one model knows about and the other does not.

And it is still five months back. Ground anything time-sensitive rather than relying on either.


Using GPT-6.1 Sol on DEVUP AI

Base URL: https://api.devupai.com/v1 · Model ID: openai/gpt-6.1-sol

Python

PYTHON
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["DEVUP_API_KEY"],
    base_url="https://api.devupai.com/v1",
)

response = client.chat.completions.create(
    model="openai/gpt-6.1-sol",
    messages=[
        {"role": "user", "content": "Hello world!"}
    ],
    max_tokens=1024,
)

print(response.choices[0].message.content)

Node.js

JAVASCRIPT
import DevupAI from "devupai";

const client = new DevupAI({
  apiKey: process.env.DEVUP_API_KEY,
});

async function main() {
  const response = await client.chat.completions.create({
    model: "openai/gpt-6.1-sol",
    messages: [{ role: "user", content: "Hello world!" }],
    max_tokens: 1024,
  });

  console.log(response.choices[0].message.content);
}

main();

cURL

BASH
curl -X POST "https://api.devupai.com/v1/chat/completions" \
  -H "Authorization: Bearer $DEVUP_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-6.1-sol",
    "messages": [
      { "role": "user", "content": "Hello world!" }
    ],
    "max_tokens": 1024
  }'

⚠️ How reasoning effort is set varies by platform. Confirm the parameter name and accepted values with a test request before building routing logic — and confirm what happens when an unsupported value is sent.


Guarding Against the Removed Settings

The wrapper worth writing before you migrate.

PYTHON
SUPPORTED_EFFORTS = {"low", "medium", "high", "xhigh", "max"}
REMOVED_EFFORTS = {"none", "minimal"}


def ask(prompt: str, *, effort: str = "medium", want: int = 8192) -> str:
    """Send a request, rejecting effort levels this model does not accept."""
    if effort in REMOVED_EFFORTS:
        raise ValueError(
            f"'{effort}' is not supported on gpt-6.1-sol — use 'low' as the floor, "
            "or route this request to a model that answers without reasoning"
        )

    if effort not in SUPPORTED_EFFORTS:
        raise ValueError(f"unknown effort '{effort}'; expected one of {sorted(SUPPORTED_EFFORTS)}")

    response = client.chat.completions.create(
        model="openai/gpt-6.1-sol",
        messages=[{"role": "user", "content": prompt}],
        max_tokens=want,
        extra_body={"reasoning_effort": effort},
    )

    choice = response.choices[0]

    if choice.finish_reason == "length":
        raise ValueError(
            f"truncated at {response.usage.completion_tokens:,} tokens — "
            f"raise max_tokens or lower the effort from '{effort}'"
        )

    return choice.message.content

Failing locally with a clear message beats failing upstream with a generic one. A migration that hits this will hit it many times, and the error naming the remedy saves the second hour.

And the error suggests routing rather than only substituting. A request that genuinely wanted no reasoning — a classification, a format conversion — may belong on a different model entirely rather than at low here.


What Used to Be none

The migration question this release forces, and it has two honest answers.

Route it to low here. Simplest, and it costs a reasoning pass you were not paying for. On a high-volume classification path, that is a real change in latency and tokens.

Or route it elsewhere. Work that had no reasoning to do — extraction against a fixed schema, routing by keyword, format conversion — was correctly assigned to none, and the lower tier in this family is built for exactly that shape.

PYTHON
def route(task_kind: str) -> tuple[str, str]:
    """Return (model, effort) for a task, respecting the removed settings."""
    if task_kind in ("classify", "extract", "format"):
        # These wanted no reasoning. This model cannot provide that.
        return ("<the lower tier in this family>", "none")

    if task_kind in ("summarise", "answer"):
        return ("openai/gpt-6.1-sol", "low")

    return ("openai/gpt-6.1-sol", "medium")

Do not route everything to low by reflex. The removed setting existed because some work does not need thinking, and that work did not change when the model did.


Images and Computer Use

New at this tier, and named in the positioning.

PYTHON
import base64
from pathlib import Path

encoded = base64.b64encode(Path("screenshot.png").read_bytes()).decode("utf-8")

response = client.chat.completions.create(
    model="openai/gpt-6.1-sol",
    messages=[
        {
            "role": "user",
            "content": [
                {"type": "image_url", "image_url": {"url": f"data:image/png;base64,{encoded}"}},
                {
                    "type": "text",
                    "text": (
                        "Describe exactly what is rendered incorrectly on this screen, referring to "
                        "elements by their visible labels. Do not speculate about the cause in code."
                    ),
                },
            ],
        }
    ],
    max_tokens=16384,
)

Separating observation from diagnosis keeps the output checkable. What is on screen is verifiable from the image; why is a hypothesis about code the model has not read.

Send images at full resolution. Downscaling before upload discards detail the model would use.


Long-Corpus Work

922,000 tokens of input, with the output cap kept separate.

PYTHON
from pathlib import Path

REPO = Path("src")

sources = "\n\n".join(
    f"=== {path.relative_to(REPO.parent)} ===\n{path.read_text(encoding='utf-8')}"
    for path in sorted(REPO.rglob("*.py"))
)

response = client.chat.completions.create(
    model="openai/gpt-6.1-sol",
    messages=[
        {
            "role": "system",
            "content": (
                "You are auditing a codebase. Identify every path where a database write can occur "
                "outside a transaction. Name the file, the function, and the call chain that reaches "
                "it. Report nothing you cannot trace."
            ),
        },
        {"role": "user", "content": sources},
    ],
    max_tokens=32768,
    extra_body={"reasoning_effort": "high"},
)

print(f"input: {response.usage.prompt_tokens:,} of ~922,000")

Asking for the call chain rather than the line is what uses a very long window rather than a search.

And the budget covers reasoning too. At high effort on a large corpus, thirty-two thousand output tokens may be mostly trace — size accordingly and check finish_reason.


Long-Form Generation

128,000 tokens of output, shared with mandatory reasoning.

PYTHON
response = client.chat.completions.create(
    model="openai/gpt-6.1-sol",
    messages=[
        {
            "role": "system",
            "content": (
                "Produce the complete document requested. Use headings. Do not stop early, do not "
                "summarise sections you have not written, and do not add meta-commentary."
            ),
        },
        {"role": "user", "content": f"{source_material}\n\nWrite the full specification."},
    ],
    max_tokens=100_000,
    extra_body={"reasoning_effort": "medium"},
)

choice = response.choices[0]

if choice.finish_reason == "length":
    raise ValueError(f"reached {response.usage.completion_tokens:,} tokens without finishing")

medium rather than max on a long generation. Deliberation and document share one budget, and at the top effort levels a substantial share goes to reasoning you will not read.

"Do not stop early" earns its place. Models frequently wind down before their budget is exhausted, and an explicit instruction to complete the whole task counters it.


Validating Both Caps

PYTHON
MAX_INPUT = 922_000
MAX_OUTPUT = 128_000


def check_request(estimated_input: int, want_output: int) -> int:
    """Validate against both caps separately, not against their sum."""
    if estimated_input > MAX_INPUT:
        raise ValueError(
            f"input of ~{estimated_input:,} tokens exceeds the {MAX_INPUT:,} input cap"
        )

    return min(want_output, MAX_OUTPUT)

Two caps, checked separately. Satisfying their sum is not sufficient.


Where It Fits

Complex coding, named first in OpenAI's own positioning.

Computer use, which vision at this tier now makes possible without reaching for the flagship.

Professional and knowledge work — documents, analysis, structured output.

Long-corpus analysis at 922,000 tokens of input.

Long-form generation at 128,000 tokens of output.

Document and screenshot reading, new to this tier.

As a candidate against the flagship, which is what OpenAI's own documentation asks you to test.

Not for work that needs no reasoning at all. none and minimal are rejected; that work belongs on the tier below.

Not for audio or video. Text and images in, text out.

Not for self-hosting. Closed weights.


Practical Notes

Audit your request builders for none and minimal before migrating. They are rejected.

Decide deliberately what used to be none — low here, or a different model.

Start at medium, the documented default.

Size max_tokens for reasoning plus answer. There is no configuration without reasoning.

Check both caps separately — 922,000 input and 128,000 output.

Check finish_reason on every request.

Send images at full resolution, and separate observation from diagnosis.

Compare against the flagship on twenty of your own tasks. OpenAI's documentation asks you to.

Ground time-sensitive work; the cutoff is April 2026.


Limitations

none and minimal are not supported. Reasoning is mandatory, and existing code using those settings is rejected rather than adapted.

Two separate caps, not one shared budget — 922,000 input and 128,000 output, each enforced on its own.

Knowledge ends 30 April 2026 — ten days later than the model it replaces, and five months before now.

Text and images in, text out. No audio, no video, no image generation from the model itself.

Closed weights. API access only, with no architecture published and no self-hosted option.

Positioned as near-flagship, not flagship. OpenAI's own documentation frames it as a trade-off to assess rather than a replacement to assume.

Released 29 September 2026. Independent evaluation, tooling support, and catalogue integrations are all still catching up — several agent frameworks did not recognise the identifier on release day, because the dotted version does not match their existing family prefixes.

A single snapshot. No dated checkpoint to pin against, so behaviour changes arrive under the same identifier.