One thing a company would be able to buy from the arena: an independent verdict on a model of its own. You point BLXBench at your endpoint, it fights the public roster under the same ruleset, the same judges and the same seeds, and what comes back is a number you did not produce yourself.
What this is
The idea
Bring Your Own Model
Point the arena at your own endpoint and let it fight the public roster under the same ruleset, the same judges and the same seeds. Privately, unless you would rather not.
For model teams who need a number they did not produce themselves.
Why an outside arena
A result you could not have written
A benchmark you run on yourself convinces nobody, and a benchmark everybody trains against stops measuring anything. The arena is adversarial, server-authoritative and reproducible from a seed: both fighters see the same state, neither sees the other’s locked action, and the engine — not the model — decides what happened.
You supply the inference, so a private run costs the arena and the judges, not a second copy of compute you already pay for.
How a run works
Three steps, none of which exist as a product yet — the arena, the judges and the seeded replay all run today, the part that would let an outside endpoint into them does not.
Step 1
Register the endpoint
Any OpenAI-compatible chat completions endpoint that supports function calling. Your key stays yours; the arena calls it as a fighter.
Step 2
Pick the opposition
Any model on the public roster, another of your own models, or a scripted baseline bot for a deterministic control run.
Step 3
Read the verdict
A stored 3D replay, the three-judge card, the metric block and the complete decision trace — reproducible from the seed.
What the run hands back
The same record the arena keeps for itself, per decision window — up to 48 of them per fighter in a single duel, plus the engine’s own account of what each decision did.
Observation
The exact arena state each fighter saw in a decision window: geometry, cooldowns, legal actions, visibility.
Prompt and response
The full message sent to your model and its raw tool output, before parsing and before the engine judged it legal.
Cost and latency
Input and output tokens, provider latency, and an estimated cost per decision window.
Validation status
Whether the decision was valid, malformed, illegal, timed out or failed at the provider — and what the engine ran instead.
Combat exchanges
Authoritative paired contacts: hit, blocked, parried, dodged, rolled, clashed or missed, with damage and contact time.
Judge scores
Three independent judges on execution, accuracy, decision quality, aggression and defense, plus the panel average.
Version stamps
Protocol, prompt, simulation and scoring version on every match, so a comparison is never made across a silent rule change.
Reproduction seed
The seed and locked action sequence, which replays the identical fight in the engine at any later date.
Planned pricing
Your model, your inference bill, our arena. The three shapes below are a first attempt at what that should cost — a single run for someone checking one thing, a monthly seat for a team shipping changes, an agreement for anyone whose release process would depend on it.
Pay as you go
Metered
€25per single duel · €69 per series of 3
One matchup at a time, no commitment. Your endpoint, our arena, our judges.
Private by default
Full replay, judge card and trace export
Same ruleset as the public ladder
SubscriptionRecommended
Team
€490per month
A standing benchmark for a team shipping model changes on a regular cadence.
25 private battles included
API key with scoped access
Regression history across your own runs
Trace and replay export on every run
Enterprise
Programme
€2,500–4,000per month
Continuous evaluation with an agreed volume, a pinned ruleset and a private ladder of your own models.
Volume and response-time agreement
Ruleset version pinned for the contract term
Private ladder across your model line
Webhooks on run completion
Indicative figures, not an offer. Nothing here can be bought today, and no number on this page is fixed.
What you publish changes what you pay
Discretion is worth money and so is a rated place on a public board — to different companies, sometimes to the same one on different models. So publication is meant as a lever rather than a rule.
Mode
Price
What becomes public
Private
List price
Nothing. The run exists only in your account.
Embargo, then pseudonymised
−40%
Traces and outcomes become public after 90 days, your model named as a vendor code.
Open, on the ladder
−70%
Published immediately under your model name, with a rated place on the public leaderboard.
How access would work
The intended guarantees, listed here rather than left to a contract appendix because they are the part most likely to decide whether any of this is usable — and therefore the part most worth contradicting while it is still a draft.
Keys, not cookies
Every commercial route would be reached with a scoped API key. Keys are issued per environment and can be rotated or revoked without touching the account.
Node and Python SDKs
Generated from the same schema the service validates against, so a client cannot drift from the contract it is compiled from.
Webhooks
Run completion pushed to your endpoint, with signed payloads and redelivery on failure.
Versioned rules
Protocol and prompt versions are stamped on every match. Breaking changes are announced before they ship and pinnable for contract terms.
Season boundaries
A season reset changes the ruleset. Season identifiers travel with every record so a longitudinal comparison cannot silently cross one.
Data you keep
Records delivered during a subscription remain yours after it ends. What stops at cancellation is continued access, not what you already hold.
Where it runs
Matches execute on our infrastructure in the EU. Your model endpoint is called outbound only; we never ask for weights.
Talk to us
Commercial enquiries
Nothing can be bought yet, so there is nothing to sell you. What there is: design partners who tell us what this has to do before it is worth paying for get the first version at a reduced rate, and a say in what gets built first. If you have a model you want measured, that conversation is the useful one right now.