Axon 2 · now in general availability

A model that shows its work.

Every answer arrives with the steps, the sources and the things it chose not to do — so you can check it before your customers do.

Read the evals
1M token context38 languagesSOC 2 · HIPAA · EU hosting
axon · session 8f21c
IB

Q3 churn rose to 4.1%, driven almost entirely by self-serve accounts under 12 seats. Enterprise churn was flat at 0.6%.
arr_by_cohort.csvchurn_q3.sql+4 more
Try
Capabilities

Three things it does differently

Not a longer list of features — three design decisions that change what you can safely put it in front of.

0 900K

Long context that stays sharp

A million tokens is only useful if attention holds at the far end. Axon is trained and evaluated on retrieval at depth, not on the first page.

sql code docs

Tools it decides to use

It writes and runs code, queries your warehouse and reads your docs — and tells you which of those it skipped, and why.

re-run from here

A trace you can audit

Every answer ships with the steps that produced it. Disagree with step three and you can re-run from there.

Benchmarks

Measured where it actually breaks

Retrieval at 900k tokens of context — the case where every model's headline number stops being true.

Deep-context retrieval accuracy

Needle-in-haystack at 900k tokens · 2,000 trials per model · higher is better
Axon 2 91.4%
Model H 87.9%
Model R 84.2%
Model K 79.6%
Open baseline 68.3%

Illustrative figures for a design template — not a published evaluation. Competitor names are placeholders. Run your own evals before choosing a model.

Deploy

Runs where your data already lives

Ask it something hard.

$5 of credit on signup · No card required