Plan a harness change as a system migration
A harness switch can change agent behavior, retained state, controls, and consumption, so evaluate the full system configuration.
Super Genius Labs Editorial · Oct 2, 2026 · 5 min readOpenAI’s “automated research intern” is an internal measurement claim, not a portable productivity benchmark
OpenAI’s research-intern milestone separates agent runtime, spending, activity, and research progress into distinct measurement layers.
Super Genius Labs Editorial · Sep 8, 2026 · 4 min readBenchmark improvement is not alignment transfer. A review gate for automated alignment research
A review gate separates benchmark gains in automated alignment research from evidence that a method will transfer to another model or setting.
Super Genius Labs Editorial · Sep 1, 2026 · 5 min readFramework-agnostic agent evaluation still has an instrumentation contract
A six-part compatibility record shows whether an agent deployment supplies the telemetry that framework-agnostic evaluation expects.
Super Genius Labs Editorial · Aug 28, 2026 · 4 min readWhat a publication-date filter can prove
Date filters can enforce retrieval eligibility, but historical claims need a separate check of what each result contained at the cutoff.
Super Genius Labs Editorial · Aug 25, 2026 · 4 min readThe FDA’s GenAI-device paper is a question set, not a compliance checklist
FDA is considering a competency-based evaluation model for generative-AI medical devices. Its discussion paper offers diligence questions about the finished device, intended use, clinical confirmation, and postmarket monitoring, not adopted requirements.
Super Genius Labs Editorial · Aug 23, 2026 · 4 min readA model score is sometimes a system score. An attribution sheet for multi-model agents
A named model may sit inside a routed, multi-agent harness. Record the models, routing, tools, benchmark configuration, access tier, and evidence owner behind the score.
Super Genius Labs Editorial · Aug 8, 2026 · 3 min readOpus 5 effort belongs in the release configuration
Version model ID, thinking state, effort, output budget, endpoint, and fallback policy together, then evaluate that tuple across quality, completion, latency, turns, and cost.
Super Genius Labs Editorial · Aug 6, 2026 · 5 min readTreat agent evaluation sandboxes like connected production systems
A sandbox label does not describe every reachable system. Review egress, intermediaries, credentials, blast radius, kill authority, forensics, and notification before the run.
Super Genius Labs Editorial · Aug 5, 2026 · 6 min readFrom agent traces to release decisions
A production trace becomes release evidence only after scoring, review, regression, and an explicit decision. Automated evaluation contributes evidence, not proof of correctness.
Super Genius Labs Editorial · Aug 3, 2026 · 4 min read