evidence
Retrieval, measured in public
suitepipeline
MRR
1.00
Precision@1
1.00
nDCG@5
1.00
Pass rate
1.00
MRR by pipeline stage · golden set
Dense only
1.00
+ RRF fusion
1.00
+ Reranker
1.00
full pipeline. Each bar is produced by running the bundled corpus through the same runTrace function the live assistant uses — there are no recorded numbers on this page.
cases · + Reranker
Run your own query
retrieval trace · live pipeline
The adversarial suite scores materially worse than the golden set, and that gap is the honest part. How the pipeline works →