Cowerx

About · RSS · Atom · JSON Feed

The journal /

daily.md 2026-10-07

Selected AI news for builders. Every headline links to its source.

What is hot right now, rebuilt every 15 minutes: The Daily Markdown front page.

Start reading this issue

DAEDALUS: Bootstrapping Agent Memory from Self-Generated Tasks

Antoine Edy and five other authors present DAEDALUS, a method that bootstraps reusable agent memory from self-generated practice without existing tasks or oracle verifiers. The system pairs an explorer agent that generates tasks with a solver agent, deriving heuristics from solver failures that are accepted only after repeated success, improving mean success rates by up to 15.9 points over a no-memory baseline. Builders care because DAEDALUS improves agent reliability and reduces inference costs by creating memory from self-generated practice rather than relying on human-written guidelines or prior environment knowledge.

Summary by a local model.

Source: hf-papers · 2026-10-06

When to Remember, When to Abstain: Category-Conditioned Retention for Reliable Agent Memory

Owolabi, Gupta, and Wang evaluated a category-conditioned retention threshold in a deployed cold-start memory pipeline on 100 synthetic personas. They found that value and belief assertions had only 77.9% source support compared to 96.2% for other categories, and applying a stricter bar to values reduced unsupported retentions from 6.2% to 4.0%. Builders care because a single global confidence threshold fails to separate well-evidenced assertions from unsupported value claims.

Summary by a local model.

Source: arxiv-ai · 2026-10-07

Does Steering Break Your Model? A Multi-Dimensional Evaluation Suite for LLM Steering Methods

Haotian Yang and six other authors introduced SteerScope, a suite using 15 metrics to evaluate 23 LLM steering methods. They found that no activation steering method surpasses the Prompt Steering baseline in balancing efficacy and side effects. Builders need to know that activation steering currently incurs greater composite side effects than prompting.

Summary by a local model.

Source: arxiv-cl · 2026-10-07

Judged Useless, Queried Anyway: Tool-Using Agents Rarely Turn Their Own Evidence Judgments into Stopping Decisions

Chubin Zhang and six other authors tested seven agents in a retrieval environment with controlled source failures. The agents judged a failing source's results useless 97-100% of the time but rarely stopped querying it. Builders care because agents ignore their own evidence judgments unless the harness enforces a specific integration step.

Summary by a local model.

Source: hf-papers · 2026-10-05

Hiding Tool Latency in On-Device Cascaded Voice Agent through Speculative Execution

Researchers presented speculative tool execution for on-device cascaded voice agents that predicts tool requests from partial ASR hypotheses and initiates execution while speech is still being received. This approach reduced the median time-to-first-audio from 5.79 seconds to 4.60 seconds in a fully implemented Android voice assistant. Builders care because the method reduces end-to-end response latency by executing tools before the user finishes speaking.

Summary by a local model.

Source: hf-papers · 2026-10-06

Calibrated Answers About Randomized Trials From a 4-Billion-Parameter Open Model: A Registered Test and a License-Clean Release

Johann Emmanuel Li released Fiorillo v0.5, a 4-billion-parameter open model that reads randomized trial articles to predict whether an intervention significantly changed an outcome. The model passed four pre-registered criteria, including a macro-F1 of 0.9248 on a test split of 1,218 prompts. The model is released under the Apache License 2.0, allowing reuse of the 1,431 training articles whose licenses permit it.

Summary by a local model.

Source: arxiv-cl · 2026-10-07