NEWTYPE · OPEN BENCHMARK · OPEN BENCHMARK

NewtypeBench — a methodology proven in public

“NewtypeBench evaluates AI coding agents on full-stack feature shipping — like SWE-Bench, but for the modern AI-native stack and judged by real production tests.”

  • corpus 123 tasks
  • HF v0.2.2
  • license BSL
  • stack FastAPI · React · Vite

leaderboard — coming soon

PROVEN IN PUBLIC

We speak in numbers

Metrics already proven on the public NewtypeBench benchmark.

123
Public benchmark tasks (71 backend + 52 frontend)
32,885
Tests executed per evaluation run
3
Task-source OSS projects (42.7k★ / 3.6k★ / 44k★)
v0.2.2
Public dataset — Hugging Face (BSL)
30%
Frontier-model measurement — Claude Opus 4.7, 20-instance smoke (binary)
123/123
Schema validation passed

We do not hide the frontier model’s measured 30%. Even the best models score at this level — which is exactly why measurement matters.

EXAMPLE

This is what an evaluation looks like

Given a natural-language spec, the agent submits a 4-line patch, and the harness auto-grades it by execution results.

backend/app/api/threads.pyReconstructed example to illustrate the task flow
  @@ -141,6 +141,10 @@ def register_routes(router):
      return ThreadListResponse(items=threads)
+ @router.get("/threads/{thread_id}/export")
+ def export_thread(thread_id: str, db=Depends(get_db)):
+     thread = get_or_404(db, Thread, thread_id)
+     return render_markdown_export(thread)
  1. FAILBefore the patch — new-feature tests failing
  2. RUNrunning pytest · Playwright … 32,885 tests
  3. PASSFAIL_TO_PASS 940 passed ∧ PASS_TO_PASS 31,945 unbroken

See for yourself

The methodology, scoring logic, and dataset are all public. The company name appears nowhere in the benchmark — a neutrality principle.

WHY THIS STACK

Evaluating AI agents on the stack AI builders actually use

The evaluation stack is no arbitrary choice — each tool ranks #1 in usage for its category.

Stack survey evidence

TechnologyUsageSource
React82%State of JS 2024 · #1 in category
Vite78.1%State of JS 2024 · #1 in category
Tailwind62%State of CSS 2024 · #1 in category
FastAPI38%JetBrains 2024 · #1 in category

─ 16,209 Indeed postings for “fastapi react”

A unique axis in the benchmark landscape

BenchmarkWhat it measures
SWE-BenchBug fixes in mature libraries
FullStackBenchGeneral code quality
Terminal-BenchCLI tasks
NewtypeBenchModern AI-native stack × full-stack feature shipping