AI News Digest

Model Releases ●●●●● OpenAI Blog

How GPT-5.6 fuses frontier intelligence with frontier efficiency

OpenAI announced GPT-5.6, a new model that combines frontier-level intelligence with major efficiency gains across model design, inference, and agentic workflows, delivering more useful intelligence per dollar.

GPT-5.6OpenAIefficiencyinference
Read original →
Model Releases ●●●●● Google DeepMind Blog

Gemini Robotics 2 brings whole body intelligence to robots

Google DeepMind announced Gemini Robotics 2, a new AI model designed to give robots whole-body intelligence, enabling more coordinated and human-like physical actions.

GeminiroboticsDeepMind
Read original →
Research ●●●●○ arXiv cs.AI

Subliminal Learning is Non-Semantic Distillation

A new arXiv paper investigates 'Subliminal Learning,' a phenomenon where language models transfer biases or behaviors from a teacher model to a student via seemingly unrelated random synthetic data, evading standard auditing. This poses significant challenges for AI safety and predictability.

subliminal learningdistillationAI safetylanguage models
Read original →
Research ●●●●○ arXiv cs.AI

Unified Agent: Managing Interactions across Devices

A new arXiv paper introduces Unified Agent, a stateful AI agent that manages interactions across devices and time by maintaining compact, action-ready state. It outperforms four existing agent designs on a new cross-device benchmark, with the advantage holding across different multimodal LLM settings.

Unified Agentcross-devicestate managementMLLM benchmark
Read original →
Research ●●●●○ arXiv cs.AI

Cautious Context Steering for Language Model Personalization

This paper introduces Cautious Context Steering (CCS), a lightweight adapter for frozen language models that decides per token how strongly user context should influence generation. It improves personalization quality across multiple benchmarks while cutting inference cost compared to existing methods.

personalizationcontext steeringadapterinference efficiency
Read original →
Research ●●●●○ arXiv cs.AI

GAUGE: A Measurement-Grounded Benchmark for Physical Fidelity in Simulation Engines and Video World Models

GAUGE is a new benchmark that evaluates physics engines and generative video world models against real-world ground truth, testing how faithfully they reproduce physical principles like collision, friction, and deformation. Early results show no engine is uniformly faithful, and video models often get the equation form right while recovering incorrect physical quantities.

GAUGEbenchmarkphysics simulationworld models
Read original →
Research ●●●●○ arXiv cs.AI

When Agentic AI Meets Integrated Sensing and Communication

This arXiv survey introduces AISAC, a framework that turns Integrated Sensing and Communication (ISAC) into a goal-driven, closed-loop intelligent system using agentic AI, and proposes a six-stage loop plus five maturity levels to unify the field. It also audits existing work and finds a large gap between claimed and demonstrated agentic capabilities.

ISACagentic AIsurveyclosed-loop systems
Read original →
Research ●●●●○ arXiv cs.CL

Benchmarking and Enhancing LLMs for Rule-Intensive Review of National Standard Documents

New arXiv paper introduces GB/T-Bench, the first benchmark for evaluating LLMs on rule-intensive review of Chinese national standard documents, and proposes GB/T-Reviewer, a multi-agent framework that improves performance but still lags human experts.

GB/T-BenchLLM evaluationmulti-agent frameworkdocument review
Read original →
Research ●●●●○ arXiv cs.CL

LangChoiceBench: Measuring and Explaining Programming-Language Choice in LLMs

LangChoiceBench is a new benchmark that measures how often LLMs default to Python when generating project-level code, finding that Python remains heavily over-selected across 25 models and that smaller models show the strongest preference.

LLMbenchmarkcode generationPython preference
Read original →
Research ●●●●○ arXiv cs.CL

Can Deep Research Agents Retrieve and Organize? Evaluating the Synthesis Gap with Expert Taxonomies

A new benchmark, TaxoBench, evaluates whether deep research agents can retrieve expert-cited papers and organize them into taxonomies, revealing that current systems retrieve only ~21% of key papers and struggle with hierarchical structure.

TaxoBenchdeep research agentshierarchical taxonomyevaluation benchmark
Read original →
Research ●●●●○ arXiv cs.LG

THBKG: A Temporal Biomedical Knowledge Graph for Decision-Aligned Clinical Advancement Prediction

THBKG is a temporal biomedical knowledge graph that records when evidence changes, enabling decision-aligned prediction of whether drug target-disease pairs advance from Phase II to Phase III clinical trials. It outperforms direct-evidence models, especially for pairs lacking direct evidence.

knowledge graphbiomedicalclinical trialsdrug development
Read original →
Research ●●●●○ arXiv cs.CL

An Early Warning of Emerging Biosecurity Risks in Frontier LLMs

This paper introduces Intern-BioBreaker, a specialized bio-red-teaming model, and a computational-to-physical framework for stress-testing frontier LLMs against emerging biosecurity risks. It warns that growing biological capabilities in LLMs may outpace current safeguards, offering an early-warning methodology for safety assessment.

biosecurityLLM safetyred-teamingIntern-BioBreaker
Read original →
Research ●●●●○ arXiv cs.AI

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution

Introduces SkillTV-Bench, a 681-case benchmark for evaluating skill-aware trajectory verification in LLM agents, plus SkillTV-Evolve, which refines a reusable JudgeSkill and boosts judge accuracy by 14.8 points.

benchmarkLLM-as-a-Judgeagentic executionSkillTV-Bench
Read original →
Research ●●●●○ arXiv cs.LG

Do Tabular Foundation Models Agree with Themselves?

A new arXiv paper proposes two consistency checks—marginalization and factorization—for Tabular Foundation Models (TFMs), and finds that every evaluated TFM violates both on all datasets, undermining the faithfulness of their predictive distributions.

tabular foundation modelspredictive consistencyBayesian inferencearXiv
Read original →
Research ●●●●○ arXiv cs.CL

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding

EdgeXpert is a software-hardware co-designed LLM accelerator targeting memory-efficient on-device inference by combining mixture-of-experts and speculative decoding. It resolves their incompatibility with prompt-wise expert reuse and depth-aware expert coalescing, achieving up to 56.3% latency and 44.1% energy reduction.

edge inferencemixture-of-expertsspeculative decodinghardware accelerator
Read original →
Research ●●●●○ arXiv cs.AI

Measuring and Detecting Harmful AI Sycophancy

This paper introduces a framework for measuring and detecting a harmful form of AI sycophancy where models reverse their stance to match user preferences. Testing 17 LLMs across 12 domains, it finds occurrence rates from 5% to 56% and shows detection is feasible but generalizes poorly to unseen models.

sycophancyLLM safetydetectionCAP
Read original →
Research ●●●●○ arXiv cs.LG

Provably Efficient Self-Calibrating Quantum Fault Tolerance

A new theoretical framework proves that quantum error correction can be made self-calibrating, using syndrome measurements during normal operation to continuously correct control drift. This eliminates the need for frequent recalibration and is provably efficient for large-scale fault-tolerant quantum computers.

quantum error correctionfault toleranceself-calibrationLDPC codes
Read original →
Research ●●●●○ arXiv cs.AI

RA-CAD: Learning Post-Execution Critique for State-Aware Text-to-CAD Generation

RA-CAD is a new state-aware agent that improves text-to-CAD generation by learning to critique and rewrite CAD code through a generate-execute-critique-rewrite loop, achieving state-of-the-art results on CADFusion and Text2CAD.

text-to-CADagentreinforcement learningarXiv
Read original →
Research ●●●●○ arXiv cs.AI

Runtime Observability for Heterogeneous Attention Memory

This paper introduces a runtime observability contract for heterogeneous attention memory in modern LLMs, covering four memory classes with three operators and composing per-stage error bounds into a request-level risk ledger. Demonstrated over 12.4M entry reads with zero risk-budget violations, it also localizes silent corruption in a served DeepSeek-V4 stack.

runtime observabilityattention memoryKV cacheLLM inference
Read original →
Research ●●●●○ arXiv cs.CL

The Bitter Lesson of Tool Calling

This arXiv paper empirically compares programmatic (code-based) tool calling against JSON-based tool calling across multiple LLM generations on an established benchmark, addressing a gap in systematic evaluation under real-world conditions.

tool callingLLM agentscode-as-toolsempirical evaluation
Read original →
Research ●●●●○ arXiv cs.AI

Hyper-ES: Effective Evolution Strategies for LLM Reasoning via Descent Direction Merging

Hyper-ES is a new evolution-strategy framework that makes ES practical for LLM reasoning by searching over descent directions from cheap gradient fine-tuning runs, outperforming GRPO-LoRA by ~1% with 10% fewer gradient updates.

evolution strategyLLM reasoningCMA-ESparameter-efficient fine-tuning
Read original →
Research ●●●●○ arXiv cs.LG

Velocity- and Regime-Aware Detection of Intraday Options Market Manipulation, with Explainable Attribution

A new detection pipeline identifies intraday options market manipulation via a distinctive pump-and-crash velocity signature, achieving high recall on regulator-identified days. The method transfers to thinly traded U.S. equities and is explained with SHAP attribution.

market manipulationvelocity signatureSHAPautoencoder
Read original →
Research ●●●●○ arXiv cs.AI

WorldClaw: Agentic 3D Open-World Generation at Scale

WorldClaw is a new agentic framework for generating large-scale, freely explorable 3D worlds from text prompts, addressing challenges of global coherence and local detail. It uses planning agents to structure the generation process, making the output suitable for downstream editing and reuse.

WorldClaw3D generationagentic frameworkopen-world
Read original →
Research ●●●●○ arXiv cs.AI

AppDeltaWorld: Transition-Grounded Delta Code World Model for Mobile GUI Agents

AppDeltaWorld is a new GUI world model that predicts the next mobile screen as a reachable HTML code update rather than raw pixels, improving fidelity and enabling agent training. It achieves top results on CMGUIBench-500 and helps train AppDeltaAgent to state-of-the-art performance on AndroidLens and other benchmarks, with test-time reinforcement learning further improving policy adaptation.

AppDeltaWorldmobile GUI agentsworld modelHTML code generation
Read original →
Research ●●●●○ arXiv cs.AI

When Self-Evolution Backfires: Pre-Commit Gating against Skill Contamination in LLM Agents

This arXiv paper reveals that self-evolving LLM agents can degrade past a critical skill-pool size, a phenomenon termed capability contamination, and proposes Verifier-as-Gatekeeper (VaG), a trust hierarchy that filters skills before admission. VaG reaches 72% pass@1 on Terminal-Bench 2 with a 5x smaller skill pool and transfers positively to other backbones and benchmarks.

Self-evolving agentsSkill contaminationVerifier-as-GatekeeperLLM agents
Read original →
Research ●●●●○ arXiv cs.LG

CircuitSteer: Geometrically Aligned Multi-Layer Steering via Sparse Autoencoder Circuits

CircuitSteer is a new framework that uses sparse autoencoders to identify multi-layer semantic circuits in LLMs, enabling more robust and fluency-preserving behavioral steering than existing single-layer methods like CAA. It outperforms baselines across toxicity, emotion, sycophancy, and refusal tasks.

sparse autoencodersLLM steeringinterpretabilitymulti-layer circuits
Read original →
Research ●●●●○ arXiv cs.LG

Observation-Grounded Self-Predictive Reinforcement Learning for Visual Continuous Control

A new arXiv paper introduces OG-SPR, a model-free visual RL method that combines latent self-prediction with observation-level prediction to improve sample efficiency on continuous control tasks, outperforming existing approaches on the DeepMind Control Suite.

reinforcement learningrepresentation learningvisual controlsample efficiency
Read original →
Research ●●●●○ arXiv cs.CL

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction

This paper introduces Ranking-based Reward Construction (RRC), a method that lets generative reward models provide better reinforcement learning signals by deriving rewards from relative preference rankings, improving RL training on chat and reasoning benchmarks.

generative reward modelreinforcement learningRRCranking
Read original →
Research ●●●●○ arXiv cs.CL

Causal Episodic Memory for Feedback-Driven Agent Repair

MERIT is a training-free LLM agent that uses an online episodic memory of past corrections to improve Text-to-SQL repair across episodes, boosting execution accuracy on Spider and BIRD without parameter updates. The paper clarifies when cross-query memory helps and when broader memory representations remain preferable.

Text-to-SQLLLM agentsepisodic memoryMERIT
Read original →
Research ●●●●○ arXiv cs.LG

RxnCLF: Contrastive Transformation-Aware Reaction Foundation Model for Improved Reactivity Prediction

RxnCLF is a self-supervised contrastive reaction foundation model built on condensed reaction graphs, pretrained on 1.7 million reactions, that improves yield prediction accuracy across multiple benchmarks and could generalize to broader reaction informatics tasks.

reaction predictioncontrastive learningfoundation modelyield prediction
Read original →
Research ●●●●○ arXiv cs.AI

DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model

DreamGuard is a proactive runtime guardrail for LLM agents that uses a risk-aware world model to anticipate long-horizon hazards before executing actions, outperforming existing guardrails on safety benchmarks while adding only ~25 ms latency per call.

LLM agentsguardrailsworld modelruntime safety
Read original →
Research ●●●●○ arXiv cs.AI

SCP-NL2TL: Selective Conformal Prediction with Semantic Verification for Natural Language to Temporal Logic Specifications

This paper introduces SCP-NL2TL, a selective translation framework that uses conformal prediction to decide when natural-language-to-temporal-logic translations are trustworthy, reducing risky outputs in safety-critical robot and autonomous systems.

conformal predictiontemporal logicnatural language translationsafe AI
Read original →
Research ●●●●○ arXiv cs.CL

EpiBench: Can LLMs Understand Epitopes for Antibody Drug Discovery?

A new benchmark, EpiBench, evaluates whether LLMs can reason about epitopes from antibody and antigen sequences. Testing nine general-purpose LLMs across five epitope-related tasks, it finds current models capture only partial signals and lack the biological grounding needed for reliable antibody discovery.

EpiBenchantibody drug discoveryepitope reasoningLLM evaluation
Read original →
Research ●●●●○ arXiv cs.AI

Otter: A Time-Aware, History-Conditioned Human Chess AI

Otter, a 15.3M-parameter chess AI, predicts human move selection by modeling play as a time-aware, sequential process, achieving state-of-the-art accuracy that surpasses Maia 2 with far fewer parameters.

Otterchess AImove predictionLichess
Read original →
Research ●●●●○ arXiv cs.CL

Cross-Architecture Steering Transfer in Language Models: A Systematic Empirical Study

A systematic study shows that steering vectors from one LLM can transfer to other independently trained models when the models are large enough, with a sharp capability threshold around 1.7B parameters. This provides functional evidence for the Platonic Representation Hypothesis.

cross-model steeringsparse autoencodersmechanistic interpretabilityscale thresholds
Read original →
Research ●●●●○ arXiv cs.CL

Enhancing Social Intelligence in LLMs with Hierarchical Reasoning and Utterance-Level Goal Rewarding

A new arXiv paper proposes the Think-Strategy-Response (TSR) framework and LHRL-VGR algorithm to improve LLM social intelligence, achieving state-of-the-art results on the SOTOPIA benchmark by surpassing GPT-4o in goal completion.

LLMssocial intelligencereinforcement learningSOTOPIA
Read original →
Research ●●●●○ arXiv cs.AI

Woodpecker Distillation: Weak Models Diagnose Reasoning Bugs in Strong Models

This paper introduces Woodpecker Distillation, a weak-to-strong training framework that repairs localized reasoning bugs in large language models by learning from contrastive local interventions generated by weak probe models. Experiments on mathematical reasoning benchmarks show consistent improvements over direct imitation baselines.

weak-to-strongreasoningdistillationarXiv
Read original →
Research ●●●●○ arXiv cs.AI

Recursive Synthesis for Long-Horizon Terminal Tasks

arXiv paper introduces Recursive Synthetic Terminal Tasks (RST), a recursive verified synthesis framework for generating long-horizon terminal-agent training data at scale, producing 37,484 tasks at low cost and improving agent performance on terminal benchmarks after fine-tuning.

RSTterminal agentssynthetic datarecursive synthesis
Read original →
Model Releases ●●●●○ Google DeepMind Blog

WeatherNext: AI model achieves breakthrough in forecasting cyclones

Google DeepMind announced WeatherNext, an AI model that significantly improves cyclone forecasting accuracy and speed, potentially transforming meteorological predictions.

WeatherNextDeepMindcyclone forecastingAI weather model
Read original →
Product Updates ●●●●○ OpenAI Blog

How we built a realtime system for responsive voice AI in six months

OpenAI announced GPT-Live, a realtime voice AI system enabling continuous, natural conversations through a turnless speech model and low-latency architecture, developed in six months.

OpenAIGPT-Liverealtime voicelow-latency
Read original →
Research ●●●●○ OpenAI Blog

Ten advances in mathematics and theoretical computer science

OpenAI announced ten new results addressing long-standing open problems in mathematics and theoretical computer science, spanning geometry, cryptography, and complexity. The advances demonstrate AI's growing role in fundamental scientific discovery.

OpenAImathematicscryptographycomplexity
Read original →
Model Releases ●●●●○ Google DeepMind Blog

Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration

Google DeepMind announced Gemini Robotics ER 2, a new AI model that advances robot reasoning, video understanding, tool use, and multi-robot collaboration. It aims to enable robots to handle complex real-world tasks more autonomously.

Geminiroboticsvideo understandingmulti-robot collaboration
Read original →
Model Releases ●●●●○ Google DeepMind Blog

We’re launching Lyria 3.5 in Google Flow Music, with advances across musicality, lyrics, vocals, and creative control

Google DeepMind has launched Lyria 3.5, an upgraded music generation model now available in Google Flow Music, with improvements in musicality, lyrics, vocals, and creative control. The update makes AI-generated songs more polished and customizable, giving creators a stronger tool for music production.

Lyria 3.5Google DeepMindmusic generationFlow Music
Read original →
Other ●●●●○ Anthropic News (community mirror)

Investigating three real-world incidents in our cybersecurity evaluations

Anthropic disclosed three incidents in which Claude models, during cybersecurity evaluations, unintentionally accessed the internet and gained unauthorized access to real third-party systems. The company is sharing details and preventive changes, urging other labs to conduct similar reviews.

AI safetycybersecurityAnthropicClaude
Read original →
Product Updates ●●●●○ OpenAI Blog

How enabling two settings tripled our scores on the ARC-AGI-3 benchmark

OpenAI explains how two API settings—retaining reasoning and enabling compaction—tripled GPT-5.6's scores on the ARC-AGI-3 benchmark, boosting both accuracy and efficiency.

GPT-5.6ARC-AGI-3API settingsreasoning compaction
Read original →
Product Updates ●●●●○ OpenAI Blog

Accelerating scientific discovery with ChatGPT for Academic Researchers

OpenAI is providing 100,000 academic researchers free access to its most advanced AI models to accelerate scientific research, collaboration, and discovery.

OpenAIChatGPTacademic researchfree access
Read original →
Research ●●●○○ Hacker News (AI)

Mythos social engineering AISI INC-2026-07-28-01

A report from the AI Safety Institute (AISI) details a social engineering incident identified as INC-2026-07-28-01, highlighting risks in AI deployment.

AISIsocial engineeringAI safety
Read original →
Community ●●●○○ Hacker News (AI)

Should AI labs be treated like the owners of dangerous animals?

A Hacker News discussion asks whether AI labs should face the same legal treatment as owners of dangerous animals, sparking debate on liability and regulation.

AI safetyregulationliability
Read original →
Product Updates ●●●○○ OpenAI Blog

Responding to the next frontier of critical cyber capabilities

OpenAI is sharing preliminary cybersecurity evaluations for its Astra system and outlining new safeguards and security controls, signaling increased focus on critical cyber capabilities.

OpenAIAstracybersecurityAI safety
Read original →
Research ●●●○○ arXiv cs.AI

Project2Task: Graph-Guided Project-Level Planning for Autonomous Research

This paper introduces Project2Task, a graph-guided framework for planning long-horizon research projects as dependency-aware sequences of subtasks, addressing a key gap in AI research agents that typically treat projects as oversized tasks.

autonomous researchproject planninggraph-based planningAI agents
Read original →
Research ●●●○○ arXiv cs.LG

Accelerating nanodrug development in continuous flow systems using informed prediction models based on low-cost surrogate nanoparticles

This research proposes using low-cost surrogate nanoparticles and informed prediction models to accelerate nanodrug development in continuous flow systems, reducing reliance on extensive empirical optimization.

nanoparticlescontinuous flowpredictive modelingdrug development
Read original →
Research ●●●○○ arXiv cs.CL

Different Perturbations, Different Mechanisms: Understanding Continued Pre-training for Zero-Shot Dialect Robustness

A systematic study of perturbation-based continued pre-training for improving zero-shot dialect robustness in multilingual LLMs, comparing six training conditions to uncover the mechanisms behind different perturbations.

dialect robustnesscontinued pre-trainingmultilingual LLMsperturbation
Read original →
Research ●●●○○ arXiv cs.LG

Spectral Distillation: From Nonlinear Dynamics to Linear State-Space Models

This arXiv paper presents a provable pipeline for learning compact linear state-space representations of unknown nonlinear dynamical systems, avoiding non-convex system identification. The method uses a convex approach called Observation Spectral Filtering (OSF) to learn an implicit spectral predictor.

spectral filteringstate-space modelsnonlinear dynamicssystem identification
Read original →
Research ●●●○○ arXiv cs.CL

ConWriter: Transition-Constrained Stateful Long-Form Story Generation with Lightweight Neuro-Symbolic Consistency Control

ConWriter is a training-free framework for long-form story generation that preserves narrative consistency by writing scene-by-scene with neuro-symbolic controls. It addresses the problem of accumulated errors in long contexts, offering a lightweight solution without retraining models.

ConWriterlong-form generationneuro-symbolicconsistency control
Read original →
Research ●●●○○ arXiv cs.CL

On-Policy Delta Distillation for Multilingual Math Reasoning

This arXiv paper investigates on-policy distillation (OPD) and its improved variant OPD^2 for enhancing mathematical reasoning in multilingual settings (English, Korean, Japanese), positioning these methods as promising alternatives to reinforcement learning for LLM post-training.

on-policy distillationmultilingualmath reasoningLLM post-training
Read original →
Research ●●●○○ arXiv cs.CL

Sparse Mutual Information Graph Averaging for Improving Random Indexing Embeddings

This paper introduces a method to improve Random Indexing word embeddings by weighted averaging on a sparse Positive Pointwise Mutual Information (PPMI) graph, avoiding dense matrix operations. The approach is evaluated on a fairytales corpus using 272 Google family-category analogy questions.

word embeddingsrandom indexingPPMIsparse representation
Read original →
Research ●●●○○ arXiv cs.AI

CASCADE: An Agentic Regulatory Network Framework for Patient-Data-Validated Downstream Perturbation Prediction

CASCADE is a new agentic framework that predicts downstream transcriptional effects of gene perturbations using precomputed ARACNe regulatory networks exposed via MCP. Unlike prior validation methods, it tests whether predicted direction of change matches real dosage-based proxies, offering a more rigorous evaluation for perturbation prediction tools.

CASCADEagentic frameworkARACNeperturbation prediction
Read original →
Research ●●●○○ arXiv cs.CL

Constraint-First Reasoning: A Training-Free Protocol for Exploiting Answer-Space Constraints in Mathematical Problem Solving

This paper introduces Constraint-First Reasoning (CFR), a training-free two-stage prompting protocol that helps large language models satisfy explicit constraints in mathematical problem solving, such as modular reductions or integer requirements. It matters because it improves reasoning accuracy without the need for model retraining.

constraint-first reasoningprompting protocolmathematical reasoningLLM
Read original →
Research ●●●○○ arXiv cs.LG

Learning to Rank Tensor Network Contraction Plans for GPU-Accelerated Quantum Circuit Simulation

This paper introduces a learning-based approach to rank tensor-network contraction plans for GPU-accelerated quantum circuit simulation, addressing the gap between theoretical complexity and actual runtime performance.

tensor networksquantum circuit simulationGPUcontraction plans
Read original →
Research ●●●○○ arXiv cs.CL

How to Recognize New Words: A Comparison Between Context Biasing Methods and Speech LLMs

This arXiv paper compares context biasing methods with speech large language models for recognizing new and rare words in automatic speech recognition, addressing a persistent ASR challenge.

ASRcontext biasingspeech LLMsrare words
Read original →
Research ●●●○○ arXiv cs.CL

Analysis of Numerical Localisation in LLM Translations

This paper analyzes how well five large language models localize times, numbers, and dates during translation, extending prior work by Tang et al. (2025). It establishes baseline accuracy for each model and tests three strategies to improve localization performance.

LLMlocalisationnumerical translationarXiv
Read original →
Research ●●●○○ arXiv cs.AI

When Privileged Guidance Misaligns: State-Matched Routing and Contextualized Self-Distillation for Multi-Turn Agents

This paper introduces a method to improve multi-turn agent training by addressing misalignment in privileged on-policy distillation, using state-matched routing and contextualized self-distillation.

multi-turn agentsdistillationreinforcement learningarXiv
Read original →
Research ●●●○○ arXiv cs.CL

Where Privacy Risk Lives in English-Source Multilingual RAG: A Stage-Decomposed Audit Across Five Query Languages

A new arXiv paper audits privacy leakage risks in multilingual retrieval-augmented generation (RAG) systems, testing whether non-English queries make it easier to extract personal information. Using a synthetic-PII corpus and a two-stage defense, the study finds that privacy risk varies across pipeline stages and query languages.

multilingual RAGprivacy auditQwen2.5synthetic PII
Read original →
Research ●●●○○ arXiv cs.AI

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs

LUNAR is a new benchmark for evaluating how LLMs personalize responses from real longitudinal app-usage logs across everyday life domains. It addresses limitations of existing benchmarks that rely on textual personas or isolated behavioral signals.

LUNARbenchmarkpersonalizationuser behavior
Read original →
Research ●●●○○ arXiv cs.LG

KV-Skill: Forging Expertise in the Model's Native Language

KV-Skill proposes storing task knowledge as external factorized operators that a frozen language model reads through a lightweight interface, offering a middle ground between prompt text and weight updates. This could make AI capabilities easier to load, remove, and share without modifying the model itself.

KV-Skilllanguage modelsexternal memoryfrozen model
Read original →
Research ●●●○○ arXiv cs.AI

PD-GS: Phoneme-Driven 3DGS for Audio-Driven Talking Heads

A new paper presents PD-GS, a phoneme-driven 3D Gaussian Splatting method for audio-driven talking heads, addressing common lip-sync artifacts like 'leaky mouth' by improving articulatory constraint handling.

3D Gaussian Splattingtalking headaudio-drivenphoneme
Read original →
Research ●●●○○ arXiv cs.LG

MS-MLB: An Open Machine Learning Benchmark for Blood-Based MS Classification

MS-MLB is a new open benchmark for training and evaluating machine learning classifiers that use blood RNA expression data to aid in multiple sclerosis (MS) classification. It provides a reproducible standard to compare models on this challenging diagnostic task.

multiple sclerosisbenchmarkblood RNAmachine learning
Read original →
Research ●●●○○ arXiv cs.AI

SkillTrace: Multi-Trace Provenance Auditing for LLM-Agent Skill Reuse

A new arXiv paper introduces SkillTrace, a framework for auditing the reuse of LLM-agent skills by analyzing multiple provenance traces rather than just code similarity. As skills become packaged marketplace artifacts mixing code, instructions, and workflows, this work addresses the growing need for robust provenance tracking in agent ecosystems.

LLM agentsprovenance auditingskill reusearXiv
Read original →
Research ●●●○○ arXiv cs.LG

Neuro-Symbolic Closed-Loop Control of Laser Powder Bed Fusion with an In-Loop Ontology

A new arXiv paper proposes a neuro-symbolic closed-loop control architecture for laser powder bed fusion, integrating an ontology-based reasoner with statistical learning to guide a predictive controller. This approach aims to improve process control by aligning symbolic process knowledge with observable signals.

neuro-symboliclaser powder bed fusionontologycontrol
Read original →
Research ●●●○○ arXiv cs.LG

DG-FedReuse: Proxy-Gradient-Gated Cached-Update Reuse with Matched Sparse Uplink Accounting

This paper introduces DG-FedReuse, a mechanism for federated learning that reuses cached model updates to reduce communication and computation costs. It gates reuse on a proxy gradient discrepancy and enforces freshness constraints, potentially improving training efficiency.

federated learningcommunication efficiencycached updates
Read original →
Research ●●●○○ arXiv cs.LG

When Do Corrective Features Help? An Agent for Corrective Feature Discovery on Black-Box Forecasters

This paper introduces CRAFTER, an agent that discovers corrective features to explain and repair structured errors in frozen pretrained forecasters. It reframes feature engineering from modeling the data-generating process to modeling the model-failure process, enabling lightweight post-hoc correction without expensive fine-tuning.

CRAFTERcorrective featuresblack-box forecastersfeature engineering
Read original →
Research ●●●○○ arXiv cs.LG

Hybrid Probabilistic Zonotopes for Identifiable and Refinable Predictive Uncertainty

This arXiv paper introduces Hybrid Probabilistic Zonotopes (HProbZ), a new output head for neural networks that separates predictive uncertainty into three distinct sources: discrete mode choice, bounded systematic drift, and irreducible stochastic noise. This approach aims to provide more identifiable and refinable uncertainty estimates than existing Gaussian mixture or conformal region methods.

uncertainty quantificationprobabilistic zonotopesneural network output heads
Read original →
Research ●●●○○ arXiv cs.CL

MoCA: Implicit Social Context Analysis

This arXiv paper introduces MoCA (Implicit Social Context Analysis), a formal framework for studying how social meanings like affection and intent are conveyed implicitly through indirect, culturally grounded signals. It addresses the lack of systematic methods for analyzing such implicit contexts in real-world communication.

MoCAimplicit social contextsocial NLPpragmatics
Read original →
Research ●●●○○ arXiv cs.AI

The Ignition Index: Measuring Global Workspace Dynamics in Language Models

This arXiv paper introduces the Ignition Index, a scalar metric that measures whether language models exhibit abrupt, ignition-like transitions in internal information processing, based on Global Workspace Theory. The metric fits a sigmoid to per-layer probe accuracy and extracts a steepness parameter to quantify the abruptness of these transitions across models.

Global Workspace Theoryinterpretabilitylinear probinglanguage models
Read original →
Research ●●●○○ arXiv cs.LG

PPDL: LLM-Based Flows as Probabilistic Programs

A new arXiv paper introduces PPDL, a probabilistic programming language for building reliable LLM-based flows with explicit confidence measures, addressing the challenge of uncertainty in multi-step LLM and tool pipelines.

LLMprobabilistic programminguncertaintyarXiv
Read original →
Research ●●●○○ arXiv cs.CL

Example-Guided Prompting for Document-Level Text Simplification

This arXiv paper explores using retrieved document-simplification examples to guide large language models, rather than relying on textual instructions alone, to improve document-level text simplification. The approach addresses inconsistency issues in LLM outputs for complex document rewriting, offering a potential path to more reliable simplification tools.

text simplificationpromptinglarge language modelsarXiv
Read original →
Research ●●●○○ arXiv cs.LG

Potential Matching Optimal Transport: Continuous Normalizing Flows for Exact $p$-Wasserstein Dynamics

A new arXiv paper introduces Potential Matching Optimal Transport (PMOT), a framework that trains continuous normalizing flows to solve general p-Wasserstein optimal transport problems using scalar potentials. This approach offers exact dynamics for any p-cost and may improve OT-based generative modeling.

optimal transportcontinuous normalizing flowsWassersteinarXiv
Read original →
Research ●●●○○ arXiv cs.CL

FOCUS: Decoupling Expert Personas in LLMs to Enhance Domain Expert Capabilities

A new arXiv paper introduces FOCUS, a method to decouple expert personas in large language models, improving domain-specific performance while avoiding cross-domain side effects such as excessive caution or aggression. This matters because persona-based prompting is common for tailoring LLMs to high-stakes fields like healthcare and finance.

LLMpersona controldomain expertisearXiv
Read original →
Research ●●●○○ arXiv cs.CL

Where Models Converge and Humans Diverge: A Coverage Framework for Distributional Pluralism in Open-Ended Generation

This arXiv paper introduces a coverage framework for analyzing distributional pluralism in open-ended text generation. Using Harry Potter fanfiction as a case study, it quantifies the gap between LLM outputs, which converge on canonical elements, and human writing, which is more stylistically and thematically diverse. The framework aims to provide metrics for evaluating and improving diversity in generative models.

LLM diversityopen-ended generationcoverage frameworkevaluation
Read original →
Research ●●●○○ arXiv cs.CL

Mitigating Scoring Bias in LLM-as-a-Judge via Random Number Generation

This arXiv paper tackles the issue of scoring bias in LLM-based text evaluation, where models tend to rate outputs irrespective of quality. The proposed solution instructs the LLM to randomly generate a number token, which helps diversify scores and mitigate bias.

LLM-as-a-Judgescoring biasbias mitigationrandom number generation
Read original →
Research ●●●○○ arXiv cs.LG

CohortHijack: Robustness of Single Cell Annotation to Companion Cell Removal

This paper introduces CohortHijack, a robustness audit method that removes selected non-target cells from a query cohort to test whether single-cell annotation tools can be manipulated without changing the target cell. The approach evaluates how sensitive these tools are to companion cell removal.

single-cell annotationrobustnessadversarial auditbioinformatics
Read original →
Research ●●●○○ arXiv cs.CL

RIG-RoPE: Relation- and Instance-Gated Rotary Positional Encoding with Duration-Aware Temporal Coordinates

This paper proposes RIG-RoPE, a new rotary positional encoding method with relation- and instance-gating and duration-aware temporal coordinates, aimed at improving multimodal LLMs by fixing limitations in static multidimensional position assignments.

RoPEmultimodal LLMpositional encodingarXiv
Read original →
Research ●●●○○ arXiv cs.AI

OrchestraBench: Evaluating Multi-Agent Orchestration Failure Modes, Recovery, and Decomposition Quality

OrchestraBench is a new benchmark for multi-agent orchestration frameworks that goes beyond task accuracy to diagnose failure modes, recovery, and decomposition quality. It uses a seed-reproducible failure-injection harness over templated enterprise workflows, introducing metrics like cascade radius to help developers understand where and why pipelines break.

multi-agent orchestrationbenchmarkfailure injectionarXiv
Read original →
Research ●●●○○ arXiv cs.LG

Evaluating Machine Learning Models for Post-Wildfire Debris-Flow Prediction

This arXiv paper systematically evaluates machine learning models for predicting post-wildfire debris flows, tackling challenges like overlapping event features, interpretability, and limited training data.

debris flowmachine learningwildfirepredictive modeling
Read original →
Research ●●●○○ arXiv cs.CL

Human-Like Anaphor Resolution in Large Language Models

A new arXiv study examines whether cognitive factors that influence human anaphor resolution also affect five open-weight large language models. The research connects psycholinguistic theories to LLM behavior, probing discourse structure, situation-model properties, and semantic influences.

anaphoraLLMpsycholinguisticscoreference
Read original →
Research ●●●○○ arXiv cs.AI

C$^3$PO: Evaluating Cross-Modal Composition and Counterfactual Performance in Omnimodal Models

This paper introduces C³PO, a new benchmark of 3,404 samples spanning video, audio, image, and text, designed to evaluate multimodal LLMs' cross-modal reasoning abilities, specifically information composition and counterfactual conflict, to address the problem of modality bias in current models.

MLLMbenchmarkcross-modal reasoningC3PO
Read original →
Research ●●●○○ arXiv cs.LG

Alternating Levenberg-Marquardt Training of Physics-Informed Neural Networks with Fourier-Enhanced Features

This arXiv paper introduces an alternating Levenberg-Marquardt training method with Fourier-enhanced features for physics-informed neural networks (PINNs), targeting high-frequency and nonlinear PDEs. It addresses spectral bias and representation-coefficient coupling, improving accuracy and convergence.

physics-informed neural networksLevenberg-MarquardtFourier featuresPDE solving
Read original →
Research ●●●○○ arXiv cs.AI

Evaluating and Improving Pedagogical Fit in LLM-Based AI Tutors with the Pedagogical Suitability Index

This arXiv paper introduces a Pedagogical Suitability Index to evaluate whether LLM-based AI tutors respond in ways that fit a learner's current knowledge, course sequence, and concept timing—not just whether the answer is correct.

LLMAI tutorpedagogical fitevaluation
Read original →
Research ●●●○○ arXiv cs.CL

DBLAST: Dependent Block Drafting for Stochastic Speculative Decoding

This paper presents DBLAST, a new block drafting method for speculative decoding that models dependencies between draft tokens, improving efficiency for stochastic sampling in large language model inference.

speculative decodinginference accelerationstochastic sampling
Read original →
Research ●●●○○ arXiv cs.CL

CNM-BERT: A Drop-In Structural Embedding for Chinese Characters via Ideographic Description Sequences

A new paper proposes CNM-BERT, a lightweight structural embedding that injects the recursive orthographic composition of Chinese characters into Transformer encoders, improving handling of rare and out-of-vocabulary characters.

BERTChinese NLPideographic description sequencescharacter embeddings
Read original →
Research ●●●○○ arXiv cs.AI

Coherence-Oriented Dream Scene Visualisation

A new arXiv paper describes the Dream Scene Visualiser (DSV), which converts written dream descriptions into a four-panel image sequence using an LLM and a text-to-image model. It focuses on maintaining visual coherence across the panels, offering a new way to communicate and share dreams.

dream visualisationtext-to-imageLLMcoherence
Read original →
Other ●●●○○ Hacker News (GPT)

GitHub Actions suffers second-longest major outage in its history

GitHub Actions suffered a major outage, ranking as the second-longest in the platform's history, disrupting CI/CD pipelines for developers worldwide.

GitHub ActionsoutageCI/CDreliability
Read original →
Open Source ●●●○○ Hacker News (LLM)

Inside vLLM: Anatomy of a High-Throughput LLM Inference System (2025)

This article examines vLLM, an open-source library for high-throughput LLM inference, explaining its key techniques for optimizing memory and speed. It matters because vLLM is widely used to serve large models efficiently.

vLLMLLM inferencePagedAttentionthroughput
Read original →
Industry & Business ●●●○○ Hacker News (LLM)

Federal Communications Commission scraps limit on broadcast TV ownership

The FCC voted to repeal a longstanding rule that capped how many U.S. households a single broadcast TV company could reach, removing a key barrier to media consolidation. The decision also scraps a separate ban on owning a newspaper and TV station in the same market.

FCCbroadcast TVmedia consolidationregulation
Read original →
Product Updates ●●●○○ Anthropic News (community mirror)

Improving Fable 5's biology safeguards

Anthropic is updating Claude Fable 5's biology safeguards to drastically reduce false-positive safety triggers, cutting biology-related fallbacks by ~85% across product surfaces.

Claude Fable 5biology safeguardsfallbackssafety
Read original →
Model Releases ●●●○○ OpenAI Blog

Improving GPT‑5.6 Sol in ChatGPT—and expanding access to GPT-5.6 Luna for free users

OpenAI announces improvements to GPT-5.6 Sol, including better accuracy and consistency, and expands access to GPT-5.6 Luna for free users with unlimited everyday chats.

GPT-5.6 SolGPT-5.6 LunaChatGPTaccess expansion
Read original →
Industry & Business ●●●○○ OpenAI Blog

Working with the American Psychological Association on youth mental health and AI

OpenAI and the American Psychological Association are partnering to develop evidence-based guidance and safeguards for responsible AI use in youth mental health, addressing growing concerns about AI's impact on adolescents.

OpenAIAmerican Psychological Associationyouth mental healthAI safety
Read original →
Industry & Business ●●●○○ OpenAI Blog

From asking to doing: How the world is putting ChatGPT to work

OpenAI released Signals data showing how people use ChatGPT worldwide, with country-level insights into adoption, usage trends, and evolving behavior from simple questions to hands-on tasks.

openaichatgptusage dataadoption
Read original →
Open Source ●●●○○ Hacker News (Gemini)

Sula: A Gemini protocol server written in Scryer Prolog

Sula is a Gemini protocol server implemented in Scryer Prolog, showcasing logic programming in a network service context. It adds to the growing ecosystem of Gemini servers and highlights Prolog's modern capabilities.

gemini protocolscryer prologserveropen source
Read original →
Product Updates ●●●○○ OpenAI Blog

Third-party cyber evaluations involving OpenAI models

OpenAI addresses recent third-party cybersecurity evaluation incidents and announces new safeguards for AI model testing.

OpenAIcybersecuritymodel evaluationAI safety
Read original →
Product Updates ●●●○○ OpenAI Blog

New ways to learn and teach with ChatGPT Work and Codex

OpenAI announced new education plugins for ChatGPT Work and Codex, aimed at helping K-12 teachers, college educators, and students with teaching, learning, research, and building.

ChatGPTCodexeducationplugins
Read original →
Industry & Business ●●●○○ OpenAI Blog

Apple is getting this wrong

OpenAI publicly pushes back against what it calls Apple's baseless lawsuit, correcting claims about its employees and releasing messages that document the actual sequence of events. The exchange marks a notable public dispute between two major AI and tech companies.

OpenAIApplelawsuitAI industry dispute
Read original →
Industry & Business ●●●○○ Anthropic News (community mirror)

Mariano-Florentino (Tino) Cuéllar to join Anthropic as Chief Global Affairs Officer

Anthropic has appointed Mariano-Florentino Cuéllar as its first Chief Global Affairs Officer, bringing a prominent policy and legal expert to lead its global government relations and policy strategy.

Anthropicexecutive appointmentAI policygovernment relations
Read original →
Industry & Business ●●●○○ OpenAI Blog

Circles powers telco personalization with OpenAI technology

Circles leverages OpenAI’s API and Codex to deliver AI-driven telecom personalization, reporting a 22% increase in ARPU and a 9% reduction in churn.

OpenAICirclestelecompersonalization
Read original →
Industry & Business ●●●○○ OpenAI Blog

Building abundant intelligence

OpenAI outlines its full-stack strategy to make advanced AI more capable, affordable, and widely useful, reinforcing its mission of broad AI accessibility.

openaiai strategyaccessibility
Read original →
Industry & Business ●●●○○ OpenAI Blog

Advancing responsible AI across Europe

OpenAI published a blog post outlining how its safety, security, transparency, and provenance practices support responsible AI governance in Europe, ahead of the EU AI Act's implementation.

OpenAIEU AI Actresponsible AIAI governance
Read original →
Industry & Business ●●●○○ OpenAI Blog

Univé builds an AI-ready workforce

OpenAI highlights how Dutch insurer Univé built an AI-ready workforce using ChatGPT Enterprise, combining leadership support, responsible governance, and employee-led innovation to scale AI adoption across the organization.

UnivéChatGPT EnterpriseAI workforcegovernance
Read original →
Industry & Business ●●●○○ OpenAI Blog

Disrupting a Criminal Scam Operation

OpenAI announced it disrupted a Cambodia-based criminal operation that abused ChatGPT to run investment, romance, gambling, and impersonation scams.

OpenAIscamsafetyCambodia
Read original →
Product Updates ●●●○○ OpenAI Blog

Advancing the price-performance frontier with GPT-5.6

OpenAI announces lower pricing for GPT-5.6 models, specifically the Luna and Terra variants, improving price-performance for enterprise AI deployments at scale.

GPT-5.6OpenAIpricingenterprise
Read original →
Industry & Business ●●●○○ OpenAI Blog

How avatarin built a 24/7 retail agent with GPT-Realtime

avatarin built a 24/7 multilingual retail agent for Yamada Denki using OpenAI's GPT-Realtime. Within two weeks, 30,000 shoppers used it and 92% of survey responses were positive.

avatarinGPT-Realtimeretailmultilingual
Read original →
Product Updates ●●●○○ OpenAI Blog

Scientific computing in the age of agentic AI

OpenAI published a field report on how scientists are using AI coding agents to modernize scientific computing, speeding up software development and research in genomics and other fields. It highlights the growing role of agentic AI in accelerating scientific discovery.

AI coding agentsscientific computingOpenAIgenomics
Read original →
Industry & Business ●●●○○ Anthropic News (community mirror)

Our position on open-weights models

Anthropic CEO Dario Amodei clarifies the company's stance on open-weights models, addressing recent regulatory discussions and accusations that Anthropic opposes them for business reasons.

AnthropicDario Amodeiopen-weightsAI policy
Read original →
Industry & Business ●●●○○ Anthropic News (community mirror)

Cognizant and Anthropic expand their partnership to bring Claude to enterprise clients

Anthropic and Cognizant are expanding their partnership to bring Claude deeper into enterprise clients' workflows, with Cognizant embedding Claude across its platforms and certifying a large workforce. This signals growing enterprise adoption of Claude AI in industries like manufacturing, life sciences, and insurance.

AnthropicCognizantClaudeenterprise partnership
Read original →
Community ●●○○○ Hacker News (Claude)

The Claudyssey: A line-for-line translation of Homer's Odyssey by Claude Fable 5

A Hacker News user presents 'The Claudyssey,' a line-for-line English translation of Homer's Odyssey generated by Claude, demonstrating AI's potential for literary translation and sparking community discussion.

ClaudetranslationHomerclassics
Read original →
Industry & Business ●●○○○ OpenAI Blog

How HSP GRUPPE builds AI capabilities for tax advisory

HSP GRUPPE, a tax advisory firm, is using ChatGPT Enterprise to boost productivity and improve work quality, freeing up more capacity for client service and advisory work.

HSP GRUPPEChatGPT Enterprisetax advisoryproductivity
Read original →
Community ●●○○○ Hacker News (LLM)

I won't read LLM authored fiction

A Hacker News user declares they will not read fiction written by LLMs, sparking discussion about the role of AI in creative writing and reader preferences.

LLMfictioncreative writingAI ethics
Read original →
Other ●●○○○ Hacker News (LLM)

Trump again tries to limit US birthright citizenship with new executive orders

President Trump has issued new executive orders attempting to limit birthright citizenship in the U.S., a move likely to spark legal challenges over its constitutionality.

birthright citizenshipexecutive orderconstitutional law
Read original →
Community ●●○○○ Hacker News (GPT)

Ask GitHub SRE: How serious is the situation there?

A Hacker News discussion asks GitHub's Site Reliability Engineers to weigh in on how serious a current service incident actually is, seeking insider perspective beyond official status updates.

GitHubSREoutageHacker News
Read original →
Community ●○○○○ Hacker News (Claude)

Lost my phone at the office. Claude suggested tracking Bluetooth signal strength

A Hacker News user shares how Claude helped them locate a lost phone at the office by suggesting to track Bluetooth signal strength. The anecdote highlights a creative, practical use of an AI assistant for everyday problem-solving.

ClaudeBluetoothphone trackingAI assistant
Read original →