From Proposal to Verified Effect: Praxa, an Evidence-Bound Harness for Governed AI Agent Execution
Praxa tackles the gap between what an AI agent proposes and what it actually does by explicitly modeling authority, execution, external read-back, reconciliation, and promotion into a deterministic “evidence-bound” harness. The authors report four evidence lanes across unit tests, a provider-backed pilot, coordination experiments, and deployed evidence paths, concluding that the architecture is testable but that current results do not yet prove production superiority or broad user benefit.
Don’t be fooled – LLMs don’t reason
Technology Review argues that much of what users interpret as reasoning is actually shortcut behavior shaped by training and prompting, and that believing the wrong story can make deployments more fragile. The piece pushes for clearer diagnostics and evaluation that separates “seems thoughtful” from verifiable causal or logical competence.
Personalized Preference Optimization gets a geometry update with GAP-DPO
This arXiv work shows that DPO’s learning dynamics change depending on how preference-pair selection aligns with gradients tied to explicit user utility. It introduces GAP-DPO, which chooses preference pairs based on this geometry while controlling distribution shift, and reports improvements across personalized text benchmarks without treating pair selection as a harmless preprocessing step.
Microsoft’s “camouflaged” hyperscale data centers – and the community pushback
The Verge reports that a wave of hyperscale AI data centers has dramatically changed neighborhoods, even when facilities are designed to blend into the environment. In San Antonio, officials and residents raise concerns about traffic and localized impacts as projects proliferate, highlighting the friction between AI infrastructure expansion and community tolerance.
OpenAI’s Dots is enterprise-agent software that can also order dinner
With Dots, The Verge describes a new OpenAI agent platform that feels like workplace tooling with “little guys” you can coordinate, including practical task execution beyond chat. The hands-on coverage emphasizes the interface and workflow separation – what the agent is doing and what you’re doing – as OpenAI positions Dots for real productivity rather than novelty.
AI hallucinations are making entitled customers even worse
The Verge examines how customer-service interactions get harder when people cite confident AI output to override frontline workers, from allergy mistakes to other harmful refusals. The reporting shows a pattern of “ChatGPT says no” escalating conflict, turning hallucinations into social leverage rather than isolated model errors.
Core-Tail World Models: when heavy-tailed memory traces help modular agents
This arXiv paper argues that long-horizon agent memory is often heavy-tailed: a small core is repeatedly used, while rare states accumulate error under finite context. It proposes CTWM, a rank-based memory controller that keeps a summarized tail, achieving token savings and measurable improvements on synthetic and real agent benchmarks.
K-Dense BYOK – an open-source, hash-chained lab notebook for local AI research
From Hugging Face, K-Dense BYOK is positioned as a local-first research assistant that turns agent work into checkable records, not just outputs – with a Living Lab Notebook that links claims to logged actions. The project’s core promise is keeping data, code, and execution traces on the researcher’s machine so results can be reproduced long after the session.
OpenAI’s practical guide for building GPT-6 workflows – models, reasoning, tools
OpenAI lays out a hands-on guide for choosing GPT-6 models and tuning reasoning effort, then mapping prompts and skills to real tool-using workflows. The emphasis is on production readiness – coordinating tools, preparing operational workflows, and tightening the loop from proposal to execution.