ArXiV ML/AI/CV papers summary
Theme 1: Agentic Systems and the New SDLC
The transition from “LLMs as chatbots” to “LLMs as agents” is no longer a theoretical ambition; it is an engineering reality. We are witnessing the birth of “Software Engineering for Intelligence,” where the focus shifts from monolithic models to distributed, specialized systems.
- SDAD: Spec-Driven Agentic Development for the AI-Native SDLC and Skillware: A Software Ontology and Engineering Lifecycle for Persistent Behavioral Artifacts formalize the shift toward rigorous, spec-driven development.
- Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence argues that augmenting a single agent’s context is insufficient; we must organize intelligence into dynamic, evolving structures.
- PrimeAgentOrchestrator: Memory-Primed Agent Spawning for Personal AI Infrastructure and Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements address the practicalities of deployment, providing persistent memory and efficient fine-tuning for long-horizon tasks.
- MAESTRO: An LLM agent for end-to-end computational materials discovery and AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale demonstrate agents coordinating multi-scale tasks and learning through interaction with synthesized environments.
- Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills, DreamBench-SWE: A Multi-Session Memory-Hygiene Benchmark for Software Agents, and Beyond End-to-End Success: Diagnosing Failures in Long-Horizon Security LLM Agents move us toward repository-native benchmarks that measure actual skill lift and causal failure attribution.
- Testing and Evaluation of Agentic AI Systems In Military Command and Control highlights the need for “trajectory-grounded correctness” in high-stakes environments.
Theme 2: Physics-Informed and Scientific AI
We are moving away from treating data as abstract tokens and toward treating it as a reflection of physical laws. This “re-physicalization” of AI ensures models respect the underlying mechanics of the systems they model.
- Wrong-Physics Backdoors in Neural PDE Operators and Shared Physics Responses Recover Hidden Rankings in Neural Operator Libraries emphasize validating that neural operators respect physical laws rather than merely curve-fitting.
- Harmonic Torsional Diffusion for Protein-Ligand Flexible Docking and ReCurveflow: A Flow Matching Framework that Learns Curved Reaction Trajectories to Predict Transition State Geometries incorporate geometric priors to ensure physical validity.
- Hessian Interatomic Potentials without derivatives bypasses computational bottlenecks by predicting fundamental physical properties directly.
- Convergence of the Deep Galerkin Method for Finite State Mean Field Control Problems provides a rigorous mathematical foundation for using neural networks to solve high-dimensional differential equations in complex systems.
Theme 3: Trust, Auditability, and Safety
As AI systems take on high-stakes roles, the “black box” is a liability. We are building systems that are auditable by construction, tethering outputs to verifiable evidence.
- aiXamine: Unified Black-Box Evaluation of Cross-Dimensional Trade-offs in LLM Safety, Security, and Privacy and Auditable by Construction: An Ontology-Driven Framework for Trustworthy LLM Analytics in Enterprise Finance argue that safety must be integrated into the system architecture.
- Ansari: A Retrieval-Grounded Islamic AI Assistant – Architecture, Deployment, and Lessons from 140,000 Conversations, An integrated diffusion-weighted imaging processing and interpretation platform for MR-guided radiotherapy, and A Modular Agent for Reliable and Auditable Spatial Relation Verification in CT Scans demonstrate grounding through explicit, deterministic stages.
- JuryProbe: An Empirical Consensus-Risk Diagnostic for Routing Reference-Free Factuality Judge Panels to Grounded Verification and Consilience: Conformally Calibrated Communication Control for Hidden-Profile Multi-Agent Reasoning provide mechanisms for certified reasoning and uncertainty quantification.
- AEGIS: Preventing Cross-Domain Resource Abuse in MCP and ClawSentry: A Progressive Multi-Tier Security Monitor for Safeguarding Autonomous LLM Agents offer guardrails for the Model Context Protocol.
- Open-Weight Masked Introspection: Measuring What Language Models Can Report About Their Own Computation explores the current limits of model self-reporting.
Theme 4: Embodiment and World-Action Modeling
The frontier is shifting from 2D generation to World-Action Modeling, where models understand the causal relationship between actions and environmental change.
- WA-JEPA: Rethinking the Video JEPA Paradigm for World-Action Modeling in Autonomous Driving and DECOWAM: Decoupled Whole-Body World-Action Model for Legged Mobile Manipulation move toward conditional flow matching to predict future states as a function of action.
- RISE: Adaptive Imagination for World Action Models optimizes the “imagination budget” to focus computational resources on high-risk scenarios.
Theme 5: Statistical Rigor, Calibration, and Governance
We are moving toward a “statistical foundation for governance,” ensuring that models are not just rankers, but calibrated, robust, and honest estimators.
- EDGE: a closed-form directed test for the calibration of probabilistic binary classifiers and Marginally Useful: An Information-Gap Identity in Conformal Prediction address the necessity of moving beyond simple ranking to rigorous uncertainty quantification.
- Let Time Tell: Identification and Gaussian Process Estimation for Interrupted Time Series and Topological Detection of Hopf Bifurcations via Persistent Homology: A Functional Criterion from Time Series use Gaussian processes and algebraic topology to model time-series transitions.
- Wasserstein Exponential Smoothing for Distributional Time Series Forecasting and Fr'echet regression of multivariate distributions with nonparanormal transport extend regression to handle complex distributional data.
- Doubly robust inference via calibration, Double Machine Learning of Continuous Treatment Effects with Additive Instrumental Variables, and Identification and Honest Recovery from Semantic Observation Kernels: Operator Error, Coarsening, and Stability provide frameworks for robust causal inference and truth recovery.
- comprisk: A scikit-learn-compatible Python toolkit for competing-risks survival analysis bridges the gap between academic rigor and industrial application.
Theme 6: Efficiency and Evaluation-as-Search
Efficiency is a first-class design constraint, and evaluation is evolving into an adaptive, diagnostic process.
- Bern2Edge: A Neurosymbolic Compiler for Edge Deployment via Bernstein Polynomial Networks, AgentDecarbonizer: Carbon-Aware Execution for AI Agents, DAOP: Data-Aware Offloading and Predictive Pre-Calculation for Efficient MoE Inference, and STS: Efficient Sparse Attention with Speculative Token Sparsity optimize performance through hardware-model co-design and dynamic resource management.
- Evaluation-as-Search: Adaptive Discovery of Grounding Failures in Meeting Assistants and Six misconceptions about large language models: A minimal model and diagnostic taxonomy replace static benchmarks with adaptive discovery and theoretical taxonomies.
- When Failures Propagate: Causal Failure Attribution in Agentic Retrieval-Augmented Generation and AI with Authority, from Application to Silicon emphasize that combining generative AI with machine verification creates an “incorruptible referee” for safe scaling.