Skip to content
Blog tag

evaluation

Every note that touches evaluation, newest first.

9 notesLatest All notes RSS feed
Thinking

OpenAI’s “automated research intern” is an internal measurement claim, not a portable productivity benchmark

OpenAI’s research-intern milestone separates agent runtime, spending, activity, and research progress into distinct measurement layers.

Super Genius Labs Editorial · Sep 8, 2026 · 4 min read
Engineering

Benchmark improvement is not alignment transfer: a review gate for automated alignment research

A review gate separates benchmark gains in automated alignment research from evidence that a method will transfer to another model or setting.

Super Genius Labs Editorial · Sep 1, 2026 · 5 min read
Engineering

Framework-agnostic agent evaluation still has an instrumentation contract

A six-part compatibility record shows whether an agent deployment supplies the telemetry that framework-agnostic evaluation expects.

Super Genius Labs Editorial · Aug 28, 2026 · 4 min read
Thinking

What a publication-date filter can prove

Date filters can enforce retrieval eligibility, but historical claims need a separate check of what each result contained at the cutoff.

Super Genius Labs Editorial · Aug 25, 2026 · 3 min read
Thinking

The FDA’s GenAI-device paper is a question set, not a compliance checklist

FDA is considering a competency-based evaluation model for generative-AI medical devices. Its discussion paper offers diligence questions about the finished device, intended use, clinical confirmation, and postmarket monitoring—not adopted requirements.

Super Genius Labs Editorial · Aug 23, 2026 · 4 min read
Engineering

A model score is sometimes a system score: an attribution sheet for multi-model agents

A named model may sit inside a routed, multi-agent harness. Record the models, routing, tools, benchmark configuration, access tier, and evidence owner behind the score.

Super Genius Labs Editorial · Aug 8, 2026 · 3 min read
Engineering

Opus 5 effort belongs in the release configuration

Version model ID, thinking state, effort, output budget, endpoint, and fallback policy together, then evaluate that tuple across quality, completion, latency, turns, and cost.

Super Genius Labs Editorial · Aug 6, 2026 · 5 min read
Engineering

Treat agent evaluation sandboxes like connected production systems

A sandbox label does not describe every reachable system. Review egress, intermediaries, credentials, blast radius, kill authority, forensics, and notification before the run.

Super Genius Labs Editorial · Aug 5, 2026 · 6 min read
Engineering

From agent traces to release decisions

A production trace becomes release evidence only after scoring, review, regression, and an explicit decision. Automated evaluation contributes evidence, not proof of correctness.

Super Genius Labs Editorial · Aug 3, 2026 · 4 min read