Given a one hour duration task frontier models were succeeding just over 40% of the time by mid-2025 - up from 5% in late 2023. The trajectory is clear, but to trust models in a commercial context first requires careful evaluation.1
Independent
AI Evaluations
We independently evaluate what AI systems can do when deployed against real-world work. From selecting the best models to rightsizing hardware to match, we combine held-out tasks and atomic rubric scoring (developed in-house) with running real workloads on compute we control.
Real-world work
for AI models
AI model performance in deployment often falls far short of expectations. We help answer the questions every business is trying to answer; what work can we actually trust a given model with? And what does that cost? We go beyond leaderboards and hyperbolic claims, putting AI systems to work on realistic business problems.
That's how many months behind frontier capability the best open weights LLMs were at the end of last year. That gap is closing, fast. The model best suited to your task might not be served by a big cloud provider any longer.2
Costs to reach a given level of capability with LLMs are falling 40x a year. The economics of generative AI change too fast to make assumptions - disciplined measurement has become a requirement.3
Sources
- UK AISI, Frontier AI Trends Report (December 2025) - §2 Agents, Fig. 2.
- UK AISI, Frontier AI Trends Report (December 2025) - §7 Open-source models.
- Epoch AI, LLM inference price trends (Cottier et al., 12 Mar 2025).
AI Capability
Audit
Which AI model should you trust with your task? And what hardware should it run on? There's no single best answer, only the right one for the job - and the answer can change every few weeks.
- Scope
- Our audit process starts with task definition; specifying data and documents, identifying tools to be used and the processes to run.
- Scoring
- We don’t let language models mark their own work (no LLM as judge), score deterministically, grade against verifiable ground truths and run multi-epoch tests to ensure reproducible results that withstand scrutiny.
- Compute we control
- We run our tests on hardware we control - from local hardware, air-gapped in our on-premise lab to GPU-accelerated compute clusters in secure data centre facilities.
- Model + hardware selection
- Model and system architecture choices shouldn’t be made in isolation, we help you optimise hardware and model selection symbiotically.
- Independence
- Founder owned and independent, we don’t sell data to model developers, nor will you find them on our cap table. Our benchmarks are held out. There is nothing here to protect except the method and our own commercial credibility.
You get costed, task-specific model and hardware recommendations - and the evidence to support the choices - so findings can be re-calibrated when the field moves.
The audit is built for teams deploying AI on infrastructure they control - examples include defence, government, regulated industries and the ecosystem of companies that work alongside them.
Local LLMs
on Apple silicon
Benchmark #01 from Osinto is ‘🍏🗜️ The CIDER Press’. We look at what real analyst work open weights LLMs can do running air-gapped on Apple Silicon - no cloud calls, no hosted API and no LLM as judge.
The benchmark runs on consumer grade hardware any team could put on a desk and is built on Inspect - the open evaluation standard developed by the UK AI Security Institute. The proprietary task and rubric set is held-out.
First results publish shortly.