ArXiV ML/AI/CV papers summary
Theme 1: Agentic Orchestration, Reliability, and Governance
The field is rapidly transitioning from monolithic, “black-box” models to structured, multi-agent systems capable of autonomous, long-horizon task execution. As these agents move into safety-critical domains, the focus has shifted from mere predictive accuracy to process-level reliability, auditability, and institutional safety.
- Orchestration & Communication: BusMA: A Bus Communication Substrate for Multi-Agent Systems and ANASSA: An Agentic AI Orchestration Framework for Spatial Intelligence propose modular architectures that prioritize provenance and peer-to-peer communication over rigid hierarchies.
- Process Supervision & Credit: To bridge the “supervision-credit gap,” Reconciling Process Supervision with Outcome-Based Credit in Agentic Policy Optimization introduces TASPO, ensuring models learn why a step is correct. Similarly, Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair and Bioinfoysis Technical Report emphasize that reliable automation requires explicit, auditable receipts for verified states.
- Safety & Governance: Research is addressing the “enforcement gap”—where agents fail to act on detected problems—through frameworks like Why LLM Agents Collapse Without Oversight: The Enforcement Gap as the Mechanism Behind Emergence World Failures and AI Deployment Accountability Engineering: A Vision for Accountable AI in Safety-Critical Socio-Technical Systems. Furthermore, AcquireBound: Runtime Authorization for Resources Acquired by AI Agents and ActGuard: Pre-execution Action Auditing against Indirect Prompt Injection in LLM Agents provide runtime security at the agent-tool boundary.
- Institutional Risk: The Normalization of Deviance in AI Development warns of organizational drift, while Negotiating Risk Boundaries in AI for Policing Through Mixed-Stakeholder Deliberation highlights the necessity of human-in-the-loop deliberation for high-stakes deployment.
Theme 2: Mechanistic Interpretability and Representation Engineering
We are finally opening the “black box” of neural networks, moving away from “vibe-based” explanations toward formal verification and causal steering. This is the astrophysics of AI—looking inside the star to understand the fusion at its core.
- Formal Verification: Certifiably Interpretable Training of ReLU-MLPs for Boolean Tasks with Guaranteed Truth-Table Generalization and The Misery of Mechanistic Interpretability: A Formal Perspective represent a maturation toward rigorous, verifiable logic.
- Probing & Steering: PhysSAE: Mechanistic Interpretability of PINNs with Sparse Autoencoders and Transcoders Trace Visual Grounding and Hallucinations in Vision-Language Models decompose internal states into meaningful features. Even more striking is the ability to “steer” cognition, as seen in Artificial entrepreneurial cognition: Locating and causally steering an opportunity recognition dial inside large language models (LLMs).
- Geometric Signatures: Geometric Signatures of Conceptual Reorganization: A Counterfactual Embedding Framework for Detecting Scientific Revolutions uses the geometry of embeddings to quantify how scientific fields evolve, providing a quantitative lens on the history of ideas.
Theme 3: Scientific Machine Learning and Differentiable Physics
Machine learning is increasingly being used to solve the fundamental equations of the physical world. By embedding physical laws directly into architectures, researchers are creating models that are more stable, data-efficient, and physically consistent.
- Differentiable Simulators & Operators: Benchmarking Optimizers to Solve Inverse Problems with Differentiable Physics Simulators and A Variational Optimal Transport Operator on Incompressible Flow replace slow iterative solvers with neural operators. Further advancements in operator learning include Linearized PINN with pretrained nonlinear layers and ReLU Neural Network Approximation to Smooth Functional Operator: Dimensional Decay and Error Analysis.
- Structural & Kinetic Constraints: HGTO: A Unified Graph-Based Physics-Informed Formulation for Structural Topology Optimization and LiftGCN: Efficient Energy-Preserving Graph Learning via Joukowski Spectral Lifting for Finite Element Stress Prediction respect mesh structures, while Physics-Constrained Neural Surrogate for Domain Growth Prediction in Systems with Conserved Kinetics enforces conservation laws as hard constraints.
- Scientific Discovery: Interpretable Inverse Design of Metal-Organic Frameworks with Large Language Model Agents and Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment demonstrate how agents can navigate combinatorial spaces to develop proofs and discover new materials.
Theme 4: Efficiency, Scaling, and Hardware-Aware Design
As models grow and move to the edge, computational efficiency has become a primary design constraint. The field is shifting toward “hardware-aware” AI, where architecture is co-designed with the deployment environment.
- Inference-Time Optimization: LIMBO: Lifelong Inference-Time Memory and Budget Optimization for LLM Agents and AgentKV: Phase-Aware KV Eviction for Agentic LLMs treat memory as a scarce resource. Similarly, Lightning Weave: Improving the Accuracy-Efficiency Frontier of Reasoning Models through Capability Composition and T-LoopFormer: Token-Level Elastic-Depth Looped Transformers for Latent Reasoning With Dynamic Routing optimize the compute-accuracy trade-off.
- Edge Deployment: Towards Practical Precision Agriculture: Real-Time Fruit Detection and Video Analytics on Embedded Edge Hardware and LIMODENet: Attention-Free Compact Encoders for Information-Preserving Onboard Satellite Image Restoration demonstrate high performance on constrained hardware like the NVIDIA Jetson.
- Compression & Sustainability: UniRank: Unified Rank Allocation for Low-Rank LLM Compression and VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention provide new ways to shrink models, while Carbon-Aware Routing for Function Calling in Edge-Cloud LLM Systems ensures sustainability is a first-class citizen in infrastructure design.
Theme 5: Rigorous Evaluation and Human-Grounded Benchmarking
The community is moving away from “leaderboard-chasing” toward diagnostic, human-grounded, and deployment-fidelity evaluation protocols.
- Deployment-Fidelity: When Compliance Data Masquerades as Evaluation: Measurement Validity for Deployed AI Systems and Accuracy Is Not Service: A Decision-Aware Benchmark for Intermittent-Demand Forecasting warn that operational data is not a valid benchmark.
- Action-Level Reliability: Same Patient, Different Order: Action-Level Reliability of Clinical LLM Agents Under Repeated Runs and Measuring Cross-Task Behavioral Consistency in Language Model Agents highlight the need for stability reporting in real-world systems.
- Diagnostic Benchmarking: PIDS-Bench: Evaluating Prompt-Injection Detectors Under Over-Defense, Obfuscation, and Distribution Shift and MemRiskBench: Trace-Aware Risk-Preserving Evaluation for Long-Horizon LLM Agents advocate for trace-aware evaluation that avoids the pitfalls of “LLM-as-a-judge.”
- Human-Grounded Evaluation: The average-farmer illusion in language-model simulations of agricultural decisions and Before You Poll with LLMs: A Deliberative Diagnostic Framework emphasize testing whether models truly update beliefs, rather than merely mimicking human opinions.