Battle-Scarred CTO Research Series

Measure the work.
Then size the AI.

The Bates Enterprise AI Efficiency Benchmark asks a practical question: how much intelligence and infrastructure does a business task actually require? The benchmark measures useful outcomes across model size, quality, latency, memory, energy, and escalation behavior.

RESEARCH PRINCIPLE
RIGHT
SIZE
QUALITY
ENERGY
LATENCY
MEMORY
RISK
Pyrinas.co TAi enterprise Transparent AI architecture environment
PUBLIC WHITE PAPER

From Parameter Count to Human-Centered Local-First AI

Professor Timothy E. Bates' Public Research Edition connects the measured enterprise benchmark to hardware reuse, local-first deployment, environmental externalities, and a five-year human-in-the-loop AI apprenticeship strategy. It publishes the research questions, aggregate methods, findings, assumptions, limitations, and citations while deliberately withholding proprietary prompts, evaluator code, exact routing logic, customer data, production policy rules, internal endpoints, and TAi implementation mechanisms.

Public Research Edition v1.0 | August 2026. Public research artifact. The paper distinguishes measured results from modeled scenarios and does not claim that compact models or local infrastructure replace frontier systems for every workload.

10enterprise task categories
1.2B-108.6Bmodel scale evaluated
80/100provisional task-quality gate
0.1086 Whmean warm energy per routed winner
Research Snapshot

Parameter count did not predict useful enterprise work.

In the measured task-routing phase, the lowest-energy passing model for each evaluated task was a local compact or specialist model. The task-optimal winners were 8.3B parameters or smaller, despite the broader test set extending to 108.6B nominal parameters.

Right model. Right task.

Routine extraction, policy reasoning, architecture controls, summarization, secure-code work, and other enterprise tasks did not all benefit from the same model size.

Energy per correct task matters.

Watts alone are not enough. A slower model can consume more total energy even at lower instantaneous power, so efficiency must be measured against completed useful work.

Residency changes economics.

Cold-start and warm-resident behavior differed materially. Production architecture should account for model loading, reuse, concurrency, and routing - not just single-call token speed.

Escalation is a feature.

The goal is not to force every task onto a small model. The goal is to begin with the least costly trustworthy capability and escalate when quality, risk, context, or evidence requires more.

What We Measure

Useful work, not benchmark theater.

The benchmark is designed around business execution rather than a single leaderboard number.

OUTCOME

Quality and correctness

Task-level scoring, pass thresholds, refusal behavior, hallucination resistance, and human-review flags.

PERFORMANCE

Latency and throughput

Prompt and decode performance, cold versus warm execution, residency effects, and practical response time.

INFRASTRUCTURE

Memory and energy

Model footprint, accelerator memory, device power telemetry where available, Wh per task, and Wh per correct task.

ROUTING

Escalation behavior

Which model clears the task threshold at the lowest measured resource cost, and when a larger model or human should enter.

HARDWARE

Existing-compute reach

Whether useful enterprise inference can run on current or incrementally upgraded business hardware before new infrastructure is purchased.

GOVERNANCE

Evidence before authority

Results are interpreted in the context of security, policy, human approval, failure recovery, and controlled agency.

Scope note: These results are a measured research snapshot, not a claim that compact models universally outperform frontier systems. The quality gate is provisional, hardware and telemetry conditions matter, and some tasks legitimately justify heavyweight or external inference. The benchmark is designed to reveal where escalation is necessary - not to eliminate it.
Applied Through Pyrinas

Turn the benchmark into an infrastructure decision.

Pyrinas can apply the research method to a customer's actual workloads before the organization commits to a model, GPU fleet, cloud contract, or data-center-scale architecture.

Customer Work
DocumentsCodeSecurityOperationsKnowledgeDecision Support
BENCHMARK
Pyrinas AI Infrastructure Right-Sizing Assessment
Task ClassWhat work is actually being done?
Quality GateWhat counts as acceptable?
Model FitWhich intelligence clears the threshold?
Hardware FitWhat can existing compute support?
Risk BoundaryWhat must remain local or supervised?
EscalationWhat truly requires frontier capability?
Unit EconomicsWhat does a correct outcome cost?
EvidenceWhat can leadership verify?
RIGHT-SIZE
Deployment Decision
Existing PCsLocal GPUDepartment NodePrivate CloudExternal ModelHuman Review
"
An enterprise usually needs a better system before it needs a bigger model.

Professor Timothy E. Bates, Founder, CEO & CTO

Benchmark Before You Buy

What does your workload actually require?

Bring the business tasks, current infrastructure, constraints, and risk boundary. Pyrinas will help determine what can stay local, what should escalate, and what infrastructure is justified by evidence.

Measure before procurement.Reuse before replacement.Escalate only when justified.

Start with a workload conversation.

Use the main Pyrinas discovery form to request an AI Infrastructure Right-Sizing Assessment.

Request Discovery