Skip to content

evaluation

Every note that touches evaluation, newest first.

10 notesLatest All notes RSS feed

Plan a harness change as a system migration

Engineering

A harness switch can change agent behavior, retained state, controls, and consumption, so evaluate the full system configuration.

Super Genius Labs Editorial · Oct 2, 2026 · 5 min read

OpenAI’s “automated research intern” is an internal measurement claim, not a portable productivity benchmark

Thinking

OpenAI’s research-intern milestone separates agent runtime, spending, activity, and research progress into distinct measurement layers.

Super Genius Labs Editorial · Sep 8, 2026 · 4 min read

Benchmark improvement is not alignment transfer. A review gate for automated alignment research

Engineering

A review gate separates benchmark gains in automated alignment research from evidence that a method will transfer to another model or setting.

Super Genius Labs Editorial · Sep 1, 2026 · 5 min read

Framework-agnostic agent evaluation still has an instrumentation contract

Engineering

A six-part compatibility record shows whether an agent deployment supplies the telemetry that framework-agnostic evaluation expects.

Super Genius Labs Editorial · Aug 28, 2026 · 4 min read

What a publication-date filter can prove

Thinking

Date filters can enforce retrieval eligibility, but historical claims need a separate check of what each result contained at the cutoff.

Super Genius Labs Editorial · Aug 25, 2026 · 4 min read

The FDA’s GenAI-device paper is a question set, not a compliance checklist

Thinking

FDA is considering a competency-based evaluation model for generative-AI medical devices. Its discussion paper offers diligence questions about the finished device, intended use, clinical confirmation, and postmarket monitoring, not adopted requirements.

Super Genius Labs Editorial · Aug 23, 2026 · 4 min read

A model score is sometimes a system score. An attribution sheet for multi-model agents

Engineering

A named model may sit inside a routed, multi-agent harness. Record the models, routing, tools, benchmark configuration, access tier, and evidence owner behind the score.

Super Genius Labs Editorial · Aug 8, 2026 · 3 min read

Opus 5 effort belongs in the release configuration

Engineering

Version model ID, thinking state, effort, output budget, endpoint, and fallback policy together, then evaluate that tuple across quality, completion, latency, turns, and cost.

Super Genius Labs Editorial · Aug 6, 2026 · 5 min read

Treat agent evaluation sandboxes like connected production systems

Engineering

A sandbox label does not describe every reachable system. Review egress, intermediaries, credentials, blast radius, kill authority, forensics, and notification before the run.

Super Genius Labs Editorial · Aug 5, 2026 · 6 min read

From agent traces to release decisions

Engineering

A production trace becomes release evidence only after scoring, review, regression, and an explicit decision. Automated evaluation contributes evidence, not proof of correctness.

Super Genius Labs Editorial · Aug 3, 2026 · 4 min read