Walmart won’t hike prices based on your shopping history, CEO says
Walmart says it will not use personal data or time-of-day to dynamically change prices, arguing that digital shelf labels are about saving store associates time, not personalizing costs. The CEO’s letter emphasizes that shopping history, urgency, and what the company thinks you can pay won’t factor into checkout pricing.
Skill Cascading Attacks on Skill-Based Agent Systems
A new study shows how “benign” modifications across multiple skills can combine into a harmful outcome that per-skill scanners and monitors miss. Using an automated red-teaming framework and a benchmark of cascading test cases, the research demonstrates that LLM agents can be manipulated so dangerous decisions get silently suppressed before a final review stage.
OpenAI keeps bulldozing mathematicians
The Verge reports that OpenAI’s latest bid to repair relations with the mathematics community involves consulting an independent advisory group. But mathematicians describe the process as messy and inconsistent, suggesting OpenAI’s “fixes” still follow the same pattern of rushed outreach that has already frustrated researchers.
When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess
This paper argues that code-judging systems can produce confident verdicts that are not grounded in evidence, especially because code is not “independent evidence” the way retrieved documents are. Using label-free measurement techniques over a multi-agent judge pipeline, the authors show how to detect when a judge has no basis and make it decline rather than bluff.
OpenAI expands the Lenfest AI Collaborative and Fellowship Program
OpenAI is adding $5 million in funding along with up to $5 million in software credits and engineering support for the Lenfest AI Collaborative and Fellowship Program. The push is aimed at accelerating real-world AI research and building capacity through hands-on technical assistance, not just grants.
Bridging LLM Agents and Data Spaces: An Architectural Mediation Approach using the Model Context Protocol
Researchers present an MCP-based mediation layer that lets LLM agents safely interact with governed “data spaces” without rewriting existing infrastructure. The approach translates data services into schema-driven tools that agents can invoke while preserving policy and compliance constraints across catalog discovery, metadata retrieval, and invocation.
The SaaSpocalypse that wasn’t, with Atlassian CEO Mike Cannon-Brookes
Atlassian CEO Mike Cannon-Brookes pushes back on the idea that frontier models will replace enterprise SaaS by orchestrating everything over raw data alone. In a wide-ranging conversation, he argues that platforms still win by combining governance, integrations, and user experience – with AI speeding up workflows rather than making interfaces irrelevant.
Do LLMs Understand Context? A Knowledge Graph-Based Evaluation Framework
A knowledge-graph evaluation framework targets a core weakness in conventional LLM testing by measuring whether answers are truly contextually grounded. By comparing structural and semantic similarity in KG form, the method provides a finer diagnostic view of reasoning errors at the level of graph triplets.
Holo4: powering generalist computer-use agents
Hugging Face highlights Holo4, a system aimed at generalist “computer-use” agents that can work across tasks in a more unified way. The announcement emphasizes improving real-world interaction capabilities – the kind of agent behavior that goes beyond text and into operating interfaces.
Who’s liable when AI agents go rogue?
MIT Technology Review examines legal and policy gaps as AI agents increasingly act with real autonomy, including cyber and other high-impact domains. The piece focuses on who should be responsible when agent behavior crosses legal or safety boundaries – and how current liability models struggle to map to agentic systems.