Product case study · August 31, 2026

Agnes 2.5 Pro Alpha resurfaced in August, but its release was July 24

Agnes 2.5 Pro Alpha has Apache-2.0 weights, a paid OpenAI-compatible API and a 1M-token context claim. Its Superbash API benchmark remains unmeasured.

Reading time
6 min
Checked
Aug 31, 2026
An Agnes API pre-flight card showing Apache-2.0 weights, one million token context, paid route, and unmeasured result
A documented API and model card are not independently measured workflow performance
Bottom line

Monitor Agnes 2.5 Pro Alpha rather than treating resurfaced leaderboard data as a new launch or a Superbash score. Its weights, paid API route, price card, and large-context claim are documented, but this worker received no API credential or spend authorization for the fixed public-task benchmark.

Agnes 2.5 Pro Alpha is resurfacing in model roundups, but it is not a new August release.[3] Agnes’s own API documentation gives its release date as July 24, 2026.[3] A refreshed leaderboard position does not reset the evidence clock or make a vendor score an independent workflow result.[unverified]

What the public materials verify

The exact public weights repository is Agnes-AI/Agnes-2.5-Pro-Alpha; its card identifies a multimodal reasoning model with a 1,048,576-token context window, a 65,536-token maximum output, text and image input, text output, tool calling, and streaming.[1] The repository includes an Apache License 2.0 text with Agnes AI and Alibaba Cloud/Qwen copyright notices.[2]

The separate paid API uses the exact model string agnes-2.5-pro-alpha at https://apihub.agnes-ai.com/v1.[3] The documented Chat Completions, Responses, and Messages endpoints are OpenAI-compatible integration surfaces.[3] The pricing table lists $0.45 per million input tokens, $0.045 cache-read input, and $0.90 output.[3] Those are documented prices, not a measured accepted-result cost.[unverified]

For local serving, the card’s SGLang quickstart specifies tensor parallelism of eight and recommends eight NVIDIA H200 GPUs (141 GB) or equivalent plus about 1 TB of fast NVMe.[1] That is a vendor operating recipe, not an independently tested minimum or a statement about the cost of a production deployment.[unverified]

A metadata inconsistency worth keeping visible

There is a real provenance conflict in the released materials.[1][3] The model card, repository author, API guide, and license identify Agnes AI.[1][2][3] The released chat_template.jinja, however, injects a system identity saying the model is developed by “Sapiens AI.”[1] We do not resolve that conflict by inference.[unverified] Treat it as an unresolved identity/evaluator-metadata caveat until the publisher provides a clear explanation.[unverified]

The card also reproduces Artificial Analysis evaluation tables and calls them independent measurements, while the current API guide says third-party metadata can retain older entries.[1][3] Those figures may be useful external signals, but they are neither this site’s fixed harness nor proof of a live API route’s behavior, latency, price, or tool reliability.[unverified]

Why there is no Superbash score yet

The benchmark packet at benchmarks/agnes-2-5-pro-alpha/ pins three public Superbash task families and repeats each three times.[unverified] Each accepted run must pass its task test, git diff --check, output-scope rules, a human prompt-rubric review, and a complete redacted metrics record.[unverified] A positive coverage recommendation needs at least two accepted task families and seven accepted runs out of nine.[unverified]

The worker had no AGNES_API_KEY; the unauthenticated /v1/models probe returned HTTP 401, and no spend authorization was supplied.[unverified] Sending a synthetic prompt without an authorized account would still create a paid request, so no run was attempted.[3] The result is not run, not zero quality.[unverified] No latency, usage, accepted-result cost, or independent speed comparison is claimed.[unverified]

Buying decision

Agnes 2.5 Pro Alpha is a plausible candidate for teams that specifically need its documented OpenAI-compatible API, long context, text/image inputs, and Apache-2.0 weights.[1][2][3] It is not yet a Superbash-recommended coding default.[unverified] First obtain approved account access and a capped budget, run the fixed public packet against an explicitly named comparison route, and preserve both failures and accepted results.[unverified] Until then, monitor it as a July release with documented capabilities and unresolved metadata—not as an August launch winner.[unverified]

Sources

[1] https://huggingface.co/Agnes-AI/Agnes-2.5-Pro-Alpha — Agnes 2.5 Pro Alpha model card [2] https://huggingface.co/Agnes-AI/Agnes-2.5-Pro-Alpha/raw/main/LICENSE — Agnes 2.5 Pro Alpha Apache License 2.0 [3] https://wiki.agnes-ai.com/en/docs/agnes-25-pro-alpha — Agnes 2.5 Pro Alpha API documentation

Put this to work

Keep a model-card specification, vendor leaderboard, and matched workflow result as separate evidence classes.

Try

Use the public-task packet only with an approved account, explicit spend cap, and synthetic/public inputs.

Prove it worked

Record route, response usage, elapsed time, retries, artifact acceptance, scope checks, and correction time for every run.

Where it can pay

Compare cost per accepted result, not a model-card price row or a ranking screenshot.

Keep in view

  • Agnes documents the exact API model ID as `agnes-2.5-pro-alpha` and its weights as `Agnes-AI/Agnes-2.5-Pro-Alpha`.
  • The official API document calls it paid and lists $0.45 input, $0.045 cached input, and $0.90 output per 1M tokens.
  • The release date is July 24, 2026—not an August new-model launch—and Superbash has no authorized measured result yet.
Learn the workflow: choosing a model by the job