Jundong Hu and Shekar Ramachandran tested off-the-shelf small language models on four agent microtasks and found that 0 of 16 configurations met the required eligibility thresholds. The study shows that quantization to 4-bit does not move any configuration into eligibility, indicating the gap tracks model size rather than precision. Builders should place SLMs behind a baseline that meets the threshold and use the SLM only where the baseline fails.
Liangyu Teng and six other authors proposed a framework that selects LLM teams using offline profiling to measure error decorrelation and predictive divergence. The framework uses a greedy search to select a small team from a candidate pool based on a quality-complementarity objective. The framework provides computable metrics to optimize team composition beyond individual model quality.
Sumin Lee and three other authors present AgSpec, a framework that retrieves from session, workspace, and global corpora to supply missing text and adapt draft lengths for coding agents. AgSpec raises generation throughput over autoregressive decoding up to 4.76 times at batch size 16 on two repository-level multi-agent coding benchmarks. Builders care because AgSpec increases generation throughput for coding agents that repeatedly reproduce code and logs.
Yu Luo and 11 other authors introduced PoS, an inference-time framework that constructs and maintains explicit belief states for LLM agents. PoS achieved the highest overall performance on every benchmark across four tests using all three LLM backbones. Builders care because PoS detects Belief Trapping to ensure agents make meaningful progress toward their goals.
Juli Huang built a streaming-recall benchmark to disentangle retention from selection in bounded-memory evaluation. Holding access fixed, query-aware selection improves required-fact recall by 15.5 percentage points. Builders care because mixed comparisons confound retention and selection, making it difficult to isolate the source of performance gains.