Qiankai Xu proposes a framework where a frozen model acts as both solver and proposer to directly edit its own harness using run records from five diverse benchmarks. Starting from a 49-line seed, the evolved harness improved average scores by 12.64 points on out-of-distribution benchmarks. A single model can optimize its own execution code across multiple domains without requiring separate human-designed proposers.
Researchers introduced A2Z GameSpec-Bench, a benchmark of 100 long-form Game Design Documents for evaluating end-to-end game development by agents. The study found that current agents struggle to jointly satisfy interdependent requirements, though requirement-specific feedback improved GDD Fidelity by 10.9% relative to self-revision after two rounds. The benchmark provides targeted feedback to support more faithful game development.
Liangyu Teng and six other authors proposed a framework that selects LLM teams using offline profiling to measure error decorrelation and predictive divergence. The framework uses a greedy search to select a small team from a candidate pool based on a quality-complementarity objective. The framework provides computable metrics to optimize team composition beyond individual model quality.
Minki Kang and co-authors introduced Mid-Harness, which samples and verifies candidate actions before execution to improve terminal agent reliability. Using a GPT-5.6 Sol verifier with eight sampled actions raised Pass@1 from 50.00% to 68.03% on TerminalBench-Lite. Allocating test-time compute to verify actions before execution improves trajectory success at a lower token cost than generating more trajectories.
Maryam Haghifam and three other authors introduced STAM, a state-transition-aware memory framework that records state changes as new clinical entries arrive. STAM combines semantic retrieval with typed clinical relations to maintain current information in Active and superseded information in History. Builders care because STAM preserves the history needed to reconstruct treatment history and clinical trajectories.
Researchers introduced PrivMeSA, a system where a local agent controls disclosure to remote models to prevent re-identification via accumulated quasi-identifiers. The system reduces the disclosure of personal details from 98.0% to 0.2% of cases while improving task accuracy by up to 15.8 percentage points. Builders can reuse remote expertise locally without repeated exchanges by using a lesson memory that grows without additional outcome labels or parameter updates.