Routine extraction, policy reasoning, architecture controls, summarization, secure-code work, and other enterprise tasks did not all benefit from the same model size.
Measure the work.
Then size the AI.
The Bates Enterprise AI Efficiency Benchmark asks a practical question: how much intelligence and infrastructure does a business task actually require? The benchmark measures useful outcomes across model size, quality, latency, memory, energy, and escalation behavior.
RIGHTSIZE
From Parameter Count to Human-Centered Local-First AI
Professor Timothy E. Bates' Public Research Edition connects the measured enterprise benchmark to hardware reuse, local-first deployment, environmental externalities, and a five-year human-in-the-loop AI apprenticeship strategy. It publishes the research questions, aggregate methods, findings, assumptions, limitations, and citations while deliberately withholding proprietary prompts, evaluator code, exact routing logic, customer data, production policy rules, internal endpoints, and TAi implementation mechanisms.
Public Research Edition v1.0 | August 2026. Public research artifact. The paper distinguishes measured results from modeled scenarios and does not claim that compact models or local infrastructure replace frontier systems for every workload.
Parameter count did not predict useful enterprise work.
In the measured task-routing phase, the lowest-energy passing model for each evaluated task was a local compact or specialist model. The task-optimal winners were 8.3B parameters or smaller, despite the broader test set extending to 108.6B nominal parameters.
Watts alone are not enough. A slower model can consume more total energy even at lower instantaneous power, so efficiency must be measured against completed useful work.
Cold-start and warm-resident behavior differed materially. Production architecture should account for model loading, reuse, concurrency, and routing - not just single-call token speed.
The goal is not to force every task onto a small model. The goal is to begin with the least costly trustworthy capability and escalate when quality, risk, context, or evidence requires more.
Useful work, not benchmark theater.
The benchmark is designed around business execution rather than a single leaderboard number.
Quality and correctness
Task-level scoring, pass thresholds, refusal behavior, hallucination resistance, and human-review flags.
Latency and throughput
Prompt and decode performance, cold versus warm execution, residency effects, and practical response time.
Memory and energy
Model footprint, accelerator memory, device power telemetry where available, Wh per task, and Wh per correct task.
Escalation behavior
Which model clears the task threshold at the lowest measured resource cost, and when a larger model or human should enter.
Existing-compute reach
Whether useful enterprise inference can run on current or incrementally upgraded business hardware before new infrastructure is purchased.
Evidence before authority
Results are interpreted in the context of security, policy, human approval, failure recovery, and controlled agency.
Turn the benchmark into an infrastructure decision.
Pyrinas can apply the research method to a customer's actual workloads before the organization commits to a model, GPU fleet, cloud contract, or data-center-scale architecture.
An enterprise usually needs a better system before it needs a bigger model.
Professor Timothy E. Bates, Founder, CEO & CTO
What does your workload actually require?
Bring the business tasks, current infrastructure, constraints, and risk boundary. Pyrinas will help determine what can stay local, what should escalate, and what infrastructure is justified by evidence.
Start with a workload conversation.
Use the main Pyrinas discovery form to request an AI Infrastructure Right-Sizing Assessment.
Request Discovery