NEWTYPE · AGENT EVALUATION · ENTERPRISE AGENT EVALUATION

Can your agent really do the job?

We quantitatively judge an agent’s performance and quality against work actually performed by people — no domain expertise required. The judge is not human opinion but execution results.

The world already has countless agent and AI-model benchmarks

  • MMLU
  • GPQA
  • HumanEval
  • SWE-Bench
  • MMMU
  • ARC-AGI
  • Terminal-Bench
  • LiveCodeBench
  • MATH
  • WebArena
  • AIME
  • HellaSwag
  • …and hundreds more

But the agent you are building cannot be evaluated by any of them.

Enterprise work differs from general knowledge domains, and real workflows are far more complex and precise. That is why the evaluation standard must be the work your people actually performed — not a benchmark score.

WHY EVALUATION

You deployed an agent. But how do you know how good it is?

A demo is not real work capability. A HumanEval-style benchmark score is not your company’s job performance. To claim results, you must measure against the work itself.

“Only by evaluating how accurately the built agent performs against work done by people can you truly measure the outcome of an agent-building project.”

Quantified

Binary 0/1 verdicts and pass rates. Numbers instead of “it seems to work.” Numbers earn trust only when the judging criterion is objective.

NewtypeBench evaluation methodology

Domain-agnostic

The judging standard is comparison against real work artifacts, not domain expertise. The evaluator does not need to be a domain expert.

Generalization of the “answer = actually merged PR” principle

Reproducible

Run it again in a different environment six months later and get the same result. A container-pinned evaluation environment guarantees reproducibility.

Docker-based reproducible environment

HOW IT WORKS

A 4-Step Evaluation Pipeline

No LLM-as-judge. The only judges are the execution results of pytest and Playwright. That is why the evaluator needs no domain expertise.

  1. [1] Collect work history

    Secure ground truth

    Work artifacts actually produced by people become the answer key (ground truth).

  2. [2] Convert to tasks

    Reproducible packaging

    Package the work order (spec) + starting state + judging criteria into a reproducible task.

  3. [3] Run the agent

    Same conditions

    The agent performs the work under identical conditions. The answers stay hidden.

  4. [4] Auto-grade

    0 or 1

    New work passes ∧ existing work unbroken — judged by execution alone.

FAIL_TO_PASS PASS_TO_PASS → Did it accomplish the new work ∧ avoid breaking what already worked? Derived automatically by set operations — no manual labeling.

  • No exam leakage — what the agent never receives: answers, judging criteria, hints.
  • Tamper-proof grading — the harness force-injects grading criteria right before scoring. Modifying tests gets the agent nothing.
  • No subjectivity — measuring broadly collapses into “what is good code?” opinions. “Do the tests pass?” is objective.

7-stage curation pipeline

  1. Crawl
  2. Filter
  3. Distribution check
  4. Auto-extract verdict signals
  5. Instance conversion
  6. Schema validation
  7. Contamination tiering

SERVICE 01 — EVALUATION

Agent Performance Evaluation

Judged by execution against work actually performed: does it accomplish the new work, without breaking what already worked? The score is 0 or 1.

Artifact evaluation

We grade only the agent’s final output. No need to reveal your agent’s internals.

SWE-Bench-compatible patch interface

Agent-loop evaluation

Your agent connects to a real working environment and performs the work; the harness auto-extracts and grades the results.

Agent-loop interface

Version-to-version regression

Agent v N vs v N−1 relative comparison — the dashboard of your improvement loop.

Relative signal across versions

Model & agent comparison

Compare multiple candidates (models, vendors) under identical conditions to ground your adoption decision.

Evaluating 3–5 frontier models

Role boundary — we are the referee, not the coach

PartyScope of definition
We defineTask specs + grading core + environment reproducibility
You defineEverything about the agent — prompts, tools, model choice, loop strategy

─ This boundary becomes the service contract itself — the neutrality of an evaluation company.

SERVICE 02 — BUILD

Domain-Specific Agent Building

We are not evaluation-only. We build domain-specific agents ourselves and measure our own results. We practice “building designed to be evaluable” — from day one we define together “what will judge this agent’s success,” and delivery ships with a quantitative evaluation report.

Document work

Meeting-note and report summarization, document classification and translation. What people actually wrote is the grading standard.

ground truth: summaries/classifications staff actually produced

Approval workflows

Approval-document analysis, policy-violation detection, approval routing. Judged against past processing history.

ground truth: historical approval records

Data & reporting

Recurring report generation, metric aggregation, anomaly detection. Accuracy measured against existing human-made reports.

ground truth: reports people already produced

Dev & IT work

Code changes, ticket handling, first-response incident analysis. Merged PRs and closed tickets are the answers.

Direct application of the NewtypeBench methodology

Not industry expertise but “anywhere work leaves a trail” — examples of the domain-agnostic principle.

Build process — the same engine as evaluation

  1. Collect work history & convert to tasks
  2. Design & build the agent
  3. Quantitative evaluation — accuracy vs. human work
  4. Improve from the evaluation report

─ The same engine as our evaluation methodology powers quality control of the build business. Only possible when building and evaluating live in one company.

THE FLYWHEEL

Build results are proven only by evaluation

Only by measuring the agent’s accuracy against human-performed work can you speak to the ROI of a build project.

  1. ①

    Build a domain-specific agent

    Anywhere work leaves a trail

  2. ②

    Work people actually performed

    = the answers (ground truth)

  3. ③

    Quantitative evaluation

    Accuracy · regression · cost

Evaluation results feed the next improvement cycle — evaluation is the heart of the loop.

Even air-gapped, evaluation stays inside the network

The Docker-based evaluation harness runs inside the same closed network. Agents built on local LLMs are measured in-house as they are.

Isomorphic to NewtypeBench’s design philosophy

“The main purpose is the agent’s internal quality-improvement loop; the public leaderboard is a by-product”. Our build projects run their quality control on the same self-evaluation loop.

TRUST

An evaluation company’s product is, in the end, trust

Fairness-first

Tasks unfavorable to our own agents are mandatory. We promise to publish unfavorable scores too. Vendor names are excluded from the benchmark.

Contamination mitigation

Tasks are tiered public / held_out / internal_only; only post-cutoff data goes held-out, rotated by season. We control even for the chance the agent memorized the answers.

Transparency

Methodology, scoring logic, and reproduction steps are fully public. Anyone can reproduce the same results by the same procedure.

BSL source-available + public dataset

“We publish results even when they are unfavorable to us. We believe that is what qualifies an evaluation company.”─ Fairness pledge