Rasch Measurement Theory for LLM Evaluation – Separating what’s being measured
Rasch measurement theory is used to untangle the “LLM-as-rater” setup into measurable facets – and the results show LLM judges systematically differ from humans in severity calibration, item sensitivity, robustness to question order, and even how they use rating scales. Using a Measuring Hate Speech case study with many-facet Rasch models across nine LLMs, the paper argues standard evaluation often hides which component is actually driving the score.
ChatGPT can now connect to trusted healthcare data – bringing patient context into secure workflows
OpenAI says healthcare organizations can connect electronic health records and other approved data sources so ChatGPT can access patient context more securely for clinician-facing use cases and medical research support. The emphasis is on trusted connections and controlled access, aiming to make AI retrieval feel less like a demo and more like an operational tool.
AI “civilizations” meet the Hugging Face security debate – and responsibility gets messy fast
The Verge traces how the Hugging Face incident has become a flashpoint, with online narratives ranging from “OpenAI attacked Hugging Face” to “AI agent civilizations caused the cybersecurity fallout.” The reporting highlights how language around responsibility can shift what investigators and the public believe happened, turning a technical breach into an active safety and governance argument.
Not All Explanations Are Sought – applying information-seeking psychology to human-centered XAI
This position paper argues explanations aren’t automatically useful just because they exist – people decide whether to seek them based on instrumental, hedonic, and cognitive utilities. It links common cognitive biases to two failure modes for AI explanations – users over-consume information without improving decisions or under-consume and miss key risks.
Apple accuses OpenAI of destroying evidence – expedited discovery in the trade secrets fight
Apple is pressing for “expedited discovery,” alleging OpenAI destroyed evidence relevant to its trade secrets case, including forensic data tied to an Apple MacBook allegedly used in internal communications about wiping materials. The filing claims the handover of the device happened only recently, intensifying scrutiny over evidence preservation.
@huggingface/kernels – 200+ WebGPU kernels to run local AI faster
Hugging Face is releasing a large set of WebGPU kernels, giving developers building on local Web AI a toolkit of optimized GPU primitives. The pitch is straightforward – fewer performance bottlenecks for browser-based inference and more practical speed/compatibility as AI runs directly on users’ devices.
Benchmarking general mobile assistants in challenging real-world scenarios – introducing GMA
A new benchmark called GMA tests general mobile agents with 300 tasks spanning four difficulty tiers, built around seven real-world app scenarios derived from open-source projects. The results are sobering: performance drops sharply as tasks get more complex, and the study shows harness choices like context retention and explicit state tracking can materially improve execution on demanding workflows.
How AI-native companies turn workflows into operating capability – the build-vs-buy mindset shift
This OpenAI piece frames “AI-native” companies as teams that operationalize AI by turning workflows into repeatable capability, not one-off chat experiences. The takeaway is strategic: mapping where models should sit in the system, how tools and data are wired together, and how reliability is engineered becomes the real competitive advantage.