Product case study · August 31, 2026
Agnes 2.5 Pro Alpha resurfaced in August, but its release was July 24
Agnes 2.5 Pro Alpha has Apache-2.0 weights, a paid OpenAI-compatible API and a 1M-token context claim. Its Superbash API benchmark remains unmeasured.
Monitor Agnes 2.5 Pro Alpha rather than treating resurfaced leaderboard data as a new launch or a Superbash score. Its weights, paid API route, price card, and large-context claim are documented, but this worker received no API credential or spend authorization for the fixed public-task benchmark.
Agnes 2.5 Pro Alpha is resurfacing in model roundups, but it is not a new August release.[3] Agnes’s own API documentation gives its release date as July 24, 2026.[3] A refreshed leaderboard position does not reset the evidence clock or make a vendor score an independent workflow result.[unverified]
What the public materials verify
The exact public weights repository is Agnes-AI/Agnes-2.5-Pro-Alpha; its card identifies a multimodal reasoning model with a 1,048,576-token context window, a 65,536-token maximum output, text and image input, text output, tool calling, and streaming.[1] The repository includes an Apache License 2.0 text with Agnes AI and Alibaba Cloud/Qwen copyright notices.[2]
The separate paid API uses the exact model string agnes-2.5-pro-alpha at https://apihub.agnes-ai.com/v1.[3] The documented Chat Completions, Responses, and Messages endpoints are OpenAI-compatible integration surfaces.[3] The pricing table lists $0.45 per million input tokens, $0.045 cache-read input, and $0.90 output.[3] Those are documented prices, not a measured accepted-result cost.[unverified]
For local serving, the card’s SGLang quickstart specifies tensor parallelism of eight and recommends eight NVIDIA H200 GPUs (141 GB) or equivalent plus about 1 TB of fast NVMe.[1] That is a vendor operating recipe, not an independently tested minimum or a statement about the cost of a production deployment.[unverified]
A metadata inconsistency worth keeping visible
There is a real provenance conflict in the released materials.[1][3] The model card, repository author, API guide, and license identify Agnes AI.[1][2][3] The released chat_template.jinja, however, injects a system identity saying the model is developed by “Sapiens AI.”[1] We do not resolve that conflict by inference.[unverified] Treat it as an unresolved identity/evaluator-metadata caveat until the publisher provides a clear explanation.[unverified]
The card also reproduces Artificial Analysis evaluation tables and calls them independent measurements, while the current API guide says third-party metadata can retain older entries.[1][3] Those figures may be useful external signals, but they are neither this site’s fixed harness nor proof of a live API route’s behavior, latency, price, or tool reliability.[unverified]
Why there is no Superbash score yet
The benchmark packet at benchmarks/agnes-2-5-pro-alpha/ pins three public Superbash task families and repeats each three times.[unverified] Each accepted run must pass its task test, git diff --check, output-scope rules, a human prompt-rubric review, and a complete redacted metrics record.[unverified] A positive coverage recommendation needs at least two accepted task families and seven accepted runs out of nine.[unverified]
The worker had no AGNES_API_KEY; the unauthenticated /v1/models probe returned HTTP 401, and no spend authorization was supplied.[unverified] Sending a synthetic prompt without an authorized account would still create a paid request, so no run was attempted.[3] The result is not run, not zero quality.[unverified] No latency, usage, accepted-result cost, or independent speed comparison is claimed.[unverified]
Buying decision
Agnes 2.5 Pro Alpha is a plausible candidate for teams that specifically need its documented OpenAI-compatible API, long context, text/image inputs, and Apache-2.0 weights.[1][2][3] It is not yet a Superbash-recommended coding default.[unverified] First obtain approved account access and a capped budget, run the fixed public packet against an explicitly named comparison route, and preserve both failures and accepted results.[unverified] Until then, monitor it as a July release with documented capabilities and unresolved metadata—not as an August launch winner.[unverified]
Sources
[1] https://huggingface.co/Agnes-AI/Agnes-2.5-Pro-Alpha — Agnes 2.5 Pro Alpha model card [2] https://huggingface.co/Agnes-AI/Agnes-2.5-Pro-Alpha/raw/main/LICENSE — Agnes 2.5 Pro Alpha Apache License 2.0 [3] https://wiki.agnes-ai.com/en/docs/agnes-25-pro-alpha — Agnes 2.5 Pro Alpha API documentation
Put this to work
Keep a model-card specification, vendor leaderboard, and matched workflow result as separate evidence classes.
Try
Use the public-task packet only with an approved account, explicit spend cap, and synthetic/public inputs.
Prove it worked
Record route, response usage, elapsed time, retries, artifact acceptance, scope checks, and correction time for every run.
Where it can pay
Compare cost per accepted result, not a model-card price row or a ranking screenshot.
Keep in view
- Agnes documents the exact API model ID as `agnes-2.5-pro-alpha` and its weights as `Agnes-AI/Agnes-2.5-Pro-Alpha`.
- The official API document calls it paid and lists $0.45 input, $0.045 cached input, and $0.90 output per 1M tokens.
- The release date is July 24, 2026—not an August new-model launch—and Superbash has no authorized measured result yet.