Antoine Edy and five other authors present DAEDALUS, a method that bootstraps reusable agent memory from self-generated practice without existing tasks or oracle verifiers. The system pairs an explorer agent that generates tasks with a solver agent, deriving heuristics from solver failures that are accepted only after repeated success, improving mean success rates by up to 15.9 points over a no-memory baseline. Builders care because DAEDALUS improves agent reliability and reduces inference costs by creating memory from self-generated practice rather than relying on human-written guidelines or prior environment knowledge.
Owolabi, Gupta, and Wang evaluated a category-conditioned retention threshold in a deployed cold-start memory pipeline on 100 synthetic personas. They found that value and belief assertions had only 77.9% source support compared to 96.2% for other categories, and applying a stricter bar to values reduced unsupported retentions from 6.2% to 4.0%. Builders care because a single global confidence threshold fails to separate well-evidenced assertions from unsupported value claims.
Haotian Yang and six other authors introduced SteerScope, a suite using 15 metrics to evaluate 23 LLM steering methods. They found that no activation steering method surpasses the Prompt Steering baseline in balancing efficacy and side effects. Builders need to know that activation steering currently incurs greater composite side effects than prompting.
Chubin Zhang and six other authors tested seven agents in a retrieval environment with controlled source failures. The agents judged a failing source's results useless 97-100% of the time but rarely stopped querying it. Builders care because agents ignore their own evidence judgments unless the harness enforces a specific integration step.
Researchers presented speculative tool execution for on-device cascaded voice agents that predicts tool requests from partial ASR hypotheses and initiates execution while speech is still being received. This approach reduced the median time-to-first-audio from 5.79 seconds to 4.60 seconds in a fully implemented Android voice assistant. Builders care because the method reduces end-to-end response latency by executing tools before the user finishes speaking.
Johann Emmanuel Li released Fiorillo v0.5, a 4-billion-parameter open model that reads randomized trial articles to predict whether an intervention significantly changed an outcome. The model passed four pre-registered criteria, including a macro-F1 of 0.9248 on a test split of 1,218 prompts. The model is released under the Apache License 2.0, allowing reuse of the 1,431 training articles whose licenses permit it.