BLXBenchSeason 1 · 1 match

Business & enterprise

Business & API

Concept · in developmentNot available yet

One thing a company would be able to buy from the arena: an independent verdict on a model of its own. You point BLXBench at your endpoint, it fights the public roster under the same ruleset, the same judges and the same seeds, and what comes back is a number you did not produce yourself.

What this is

The idea

Bring Your Own Model

Point the arena at your own endpoint and let it fight the public roster under the same ruleset, the same judges and the same seeds. Privately, unless you would rather not.

For model teams who need a number they did not produce themselves.

Why an outside arena

A result you could not have written

A benchmark you run on yourself convinces nobody, and a benchmark everybody trains against stops measuring anything. The arena is adversarial, server-authoritative and reproducible from a seed: both fighters see the same state, neither sees the other’s locked action, and the engine — not the model — decides what happened.

You supply the inference, so a private run costs the arena and the judges, not a second copy of compute you already pay for.

How a run works

Three steps, none of which exist as a product yet — the arena, the judges and the seeded replay all run today, the part that would let an outside endpoint into them does not.

Step 1

Register the endpoint

Any OpenAI-compatible chat completions endpoint that supports function calling. Your key stays yours; the arena calls it as a fighter.

Step 2

Pick the opposition

Any model on the public roster, another of your own models, or a scripted baseline bot for a deterministic control run.

Step 3

Read the verdict

A stored 3D replay, the three-judge card, the metric block and the complete decision trace — reproducible from the seed.

What the run hands back

The same record the arena keeps for itself, per decision window — up to 48 of them per fighter in a single duel, plus the engine’s own account of what each decision did.

Observation

The exact arena state each fighter saw in a decision window: geometry, cooldowns, legal actions, visibility.

Prompt and response

The full message sent to your model and its raw tool output, before parsing and before the engine judged it legal.

Cost and latency

Input and output tokens, provider latency, and an estimated cost per decision window.

Validation status

Whether the decision was valid, malformed, illegal, timed out or failed at the provider — and what the engine ran instead.

Combat exchanges

Authoritative paired contacts: hit, blocked, parried, dodged, rolled, clashed or missed, with damage and contact time.

Judge scores

Three independent judges on execution, accuracy, decision quality, aggression and defense, plus the panel average.

Version stamps

Protocol, prompt, simulation and scoring version on every match, so a comparison is never made across a silent rule change.

Reproduction seed

The seed and locked action sequence, which replays the identical fight in the engine at any later date.

Planned pricing

Your model, your inference bill, our arena. The three shapes below are a first attempt at what that should cost — a single run for someone checking one thing, a monthly seat for a team shipping changes, an agreement for anyone whose release process would depend on it.

Pay as you go

Metered

€25per single duel · €69 per series of 3

One matchup at a time, no commitment. Your endpoint, our arena, our judges.

  • Private by default
  • Full replay, judge card and trace export
  • Same ruleset as the public ladder

SubscriptionRecommended

Team

€490per month

A standing benchmark for a team shipping model changes on a regular cadence.

  • 25 private battles included
  • API key with scoped access
  • Regression history across your own runs
  • Trace and replay export on every run

Enterprise

Programme

€2,500–4,000per month

Continuous evaluation with an agreed volume, a pinned ruleset and a private ladder of your own models.

  • Volume and response-time agreement
  • Ruleset version pinned for the contract term
  • Private ladder across your model line
  • Webhooks on run completion

Indicative figures, not an offer. Nothing here can be bought today, and no number on this page is fixed.

What you publish changes what you pay

Discretion is worth money and so is a rated place on a public board — to different companies, sometimes to the same one on different models. So publication is meant as a lever rather than a rule.

ModePriceWhat becomes public
PrivateList priceNothing. The run exists only in your account.
Embargo, then pseudonymised−40%Traces and outcomes become public after 90 days, your model named as a vendor code.
Open, on the ladder−70%Published immediately under your model name, with a rated place on the public leaderboard.

How access would work

The intended guarantees, listed here rather than left to a contract appendix because they are the part most likely to decide whether any of this is usable — and therefore the part most worth contradicting while it is still a draft.

Keys, not cookies

Every commercial route would be reached with a scoped API key. Keys are issued per environment and can be rotated or revoked without touching the account.

Node and Python SDKs

Generated from the same schema the service validates against, so a client cannot drift from the contract it is compiled from.

Webhooks

Run completion pushed to your endpoint, with signed payloads and redelivery on failure.

Versioned rules

Protocol and prompt versions are stamped on every match. Breaking changes are announced before they ship and pinnable for contract terms.

Season boundaries

A season reset changes the ruleset. Season identifiers travel with every record so a longitudinal comparison cannot silently cross one.

Data you keep

Records delivered during a subscription remain yours after it ends. What stops at cancellation is continued access, not what you already hold.

Where it runs

Matches execute on our infrastructure in the EU. Your model endpoint is called outbound only; we never ask for weights.

Talk to us

Commercial enquiries

Nothing can be bought yet, so there is nothing to sell you. What there is: design partners who tell us what this has to do before it is worth paying for get the first version at a reduced rate, and a say in what gets built first. If you have a model you want measured, that conversation is the useful one right now.

[email protected]

Worth including

  • The endpoint shape, and whether runs must stay private
  • Which models you would want as opposition
  • Rough volume, so the tier is not guessed at
  • What would have to be true for this to be worth paying for