Stable Miscalibration in Large Language Models: A Practical View of High-Confidence Errors
A new arXiv study argues that not all wrong-but-confident answers are fragile – some are “stable miscalibrations” that barely change when inputs or conditions shift slightly. By combining an output-level audit of confidence variation with an internal sensitivity probe on hidden states, the authors show prompt-induced stabilization can reduce hidden-state movement even when calibration is not improved.
Measuring Cross-Task Behavioral Consistency in Language Model Agents
Instead of judging agents only by success rate, the paper introduces a Behavioral Consistency Metric (BCM) that quantifies how consistently an agent behaves across different tasks. Using thousands of execution traces from multiple software-engineering agents, the authors find an important split: some systems are locally reproducible on a single task yet globally fragmented across tasks.
Active Perception for Embodied Disambiguation
This work targets a common robot failure mode: ambiguity caused not just by user intent but by missing physical evidence in what the robot can currently observe. The proposed active-perception framework uses a vision-language model to decide when to keep observing, request clarification, or commit to a selection, with real-robot experiments showing it can recover discriminative cues and object labels.
OpenAI funds 14 independent projects exploring new AI policy ideas to expand economic opportunity and strengthen societal resilience in the Intelligence Age
OpenAI is backing 14 separate efforts that focus on policy and governance approaches designed to help society benefit from AI while reducing downstream risks. The emphasis is on expanding economic opportunity alongside “resilience” goals, positioning policy innovation as part of the same strategy as model deployment.
The Defender’s Window
OpenAI frames cybersecurity risk and defense as a shifting window problem: attackers adapt quickly, but defenders can still gain advantage by narrowing response time and improving operational readiness. The piece highlights steps for security teams and explains how OpenAI is strengthening defenses to reduce exposure in the real world.
GPU Management (Part 2) from Dharma-AI on Hugging Face
Hugging Face’s Dharma-AI blog post digs into practical GPU management strategies aimed at making large-model workloads run more efficiently and reliably. It’s a serving- and ops-focused read that translates complex infrastructure constraints into tactics for better throughput and resource utilization.
What Flock’s defenders are missing
MIT Technology Review takes a hard look at the arguments around Flock, the company behind large-scale automatic license plate reader networks. The analysis centers on what critics and defenders tend to overlook about privacy, accountability, and the practical risks of deploying surveillance at scale.
Anthropic explains how Claude’s invisible text watermarks will work
Anthropic is detailing how it plans to apply “invisible” text watermarks to Claude outputs to meet EU transparency requirements. The report ties the approach to SynthID-Text-style watermarking that relies on detectable patterns derived from token probabilities.
OpenAI joins PORTS-Pike project, expanding community investment and supporting thousands of Southern Ohio jobs
OpenAI is joining the PORTS-Pike initiative to invest in community capacity while supporting job creation in Southern Ohio. Beyond the headline economics, the move signals how major AI companies increasingly wrap infrastructure, workforce development, and public-private partnerships into rollout strategies.
OpenAI reportedly disbanded its preparedness team
The Verge reports OpenAI dismantled its preparedness team that assessed serious model risks and worked on mitigation plans. Responsibility is said to be redistributed into existing teams such as those focused on specific areas like bio and cyber, marking another shift in how the company operationalizes safety.