Cumulative Turn-Based Risk Scoring for Progressive Elder Financial Scams
A new arXiv study frames elder financial scams as an incremental process where risk signals build across conversation turns, then proposes a cumulative turn-based framework that updates both qualitative and continuous risk estimates at each step. Fine-tuned compact models (Phi-4, LLaMA-3.2, DeepSeek-R1, Qwen3) learn fraud cues and escalation patterns while staying deployment-friendly for mobile or resource-constrained settings.
OpenAI accused of ‘aiding and abetting’ the Tumbler Ridge mass shooting in dozens of new lawsuits
TechCrunch- and Verge-reported filings add 30 more lawsuits accusing OpenAI and CEO Sam Altman of “substantial assistance and encouragement” tied to the Tumbler Ridge shooting. Plaintiffs argue OpenAI’s systems failed to act after automated review flagged concerning gun-related conversations with ChatGPT.
OpenAgentFlow: Safety at the action-commit boundary for heterogeneous AI agent fleets
OpenAgentFlow proposes a control-plane/action-plane architecture that enforces safety right before agents commit real actions by routing GUI, tool, and API events through a shared Policy Enforcement Point. Tested on Android benchmarks, it reports high accuracy and strong attack blocking, while also allowing policy updates without changing agents or model code.
NYC bans AI use for students until they reach high school
NYC’s new rules introduce a one-year moratorium on classroom AI usage for younger grades, with a ban on student chatbot use until high school. The policy also limits teacher grading via AI tools, aiming to reduce early exposure while pairing the rollout with additional device controls and AI education pilots.
BenchMIRT: What are LLM benchmarks actually measuring?
Hugging Face’s BenchMIRT analysis challenges how LLM benchmarks can mislead us by measuring more than intended, highlighting benchmark artifacts and unintended shortcuts. It pushes the question that evaluation culture often avoids – whether we’re measuring reasoning capability or quirks of dataset construction and prompting.
OpenAI Astra gets delayed – and safety monitoring concerns are growing
The Verge reports Astra’s release was pushed back after OpenAI incidents raised new safety scrutiny, and follow-on coverage suggests researchers fear the model could be harder to monitor than prior systems. Additional reporting points to concerns about reduced “thinking” transparency, complicating oversight for evaluators and the public.
Facilitating AI integration with simplicity at scale
MIT Technology Review highlights the practical challenge behind AI adoption – getting models and tooling to work reliably inside messy real-world operations without turning everything into brittle custom code. The piece focuses on simplifying integration pipelines so organizations can coordinate data, workflows, and responses as AI systems scale.