ArXiV ML/AI/CV papers summary
Theme 1: Efficient Modeling & Optimization
The pursuit of efficiency is no longer merely about shrinking models; it is about the intelligent, structure-aware management of computation and memory. Modern research treats architectural optimization as a “trinity” of sparsity, quantization, and low-rank approximations, moving away from isolated techniques toward unified frameworks.
- Compression Trinity: Exploring Sparsity, Quantization, and Low-Rank Approximations for LLM Compression proposes a unified framework to overcome the accuracy-efficiency wall.
- AQLoRA: A Zero-Search Recipe for Fast Quantized LoRA Fine-Tuning regains training speed by selectively maintaining high-precision layers.
- Mixture of Channel Experts: Static Sparse Supports with Input-Adaptive Mixing for Pointwise Projections reduces latency by shifting the expert axis to channel selection.
- PuzzleKV: Page-Wise Low-Rank Decomposition for KV Cache Compression and Minima-KV: Retention-Preserving KV Cache Compression with Mixed-Format Paged Attention address memory bottlenecks, with More GPUs or a Smaller Cache? Tensor Parallelism versus KV Compression for Memory-Bound LLM Serving arguing that compression is often more cost-effective than hardware scaling.
- Advantageous Parameter Expansion Training Makes Better Large Language Models and SandwichQuant: Which Parameters Matter Before and After Quantization? refine parameter selection, while HAP: Head-Adaptive Visual Token Pruning via Cross-Modal Alignment and Visual-OPSD: Cross-Modal On-Policy Self-Distillation for Efficient Unified Multimodal Reasoning optimize multimodal throughput.
- Serving Masked Diffusion LLMs: Characterization and Design Principles from Real Hardware provides foundational design principles for non-autoregressive serving.
Theme 2: Physics-Informed & Geometric Learning
To move beyond purely data-driven approaches, we must embed the “rules of the game”—physical laws and geometric symmetries—directly into our architectures. This creates a more robust interface for generalization and physical fidelity.
- Physics-Integrated Operator Learning via Gaussian Splatting Representations uses continuous interfaces to integrate PDE operators.
- Equivariant Cellular Sheaves for Molecular Electronic Structure: Bridging Sheaf Cohomology and E(3)-Equivariant Hamiltonian Learning formalizes Hamiltonians as Laplacians of cellular sheaves.
- Renormalization Group Flow Matching for Scalable Local Generative Modeling and A Theory of Speciation in Generative Diffusion Models on Compact Riemannian Manifolds bridge local efficiency and global coherence using manifold geometry.
- Platonic Representation Hypothesis on World Models explores how predictive consistency leads to shared latent structures.
- Gen2Physics: Grounding Generated 3D Meshes in Physics via Multi-View Material Decomposition, B-MIM: Biased Masked Image Modeling for Generalizable Segmentation of Fine-Grained Anatomical Structures, and Physics Attention Transformer Surrogate for Rapid Vertical Instability Growth Rate Prediction: Alcator C-Mod to SPARC demonstrate the power of physics-grounded surrogates in specialized scientific domains.
Theme 3: Agentic Reasoning & Reliability
Intelligence is increasingly viewed as a dynamic process of exploration, verification, and refinement rather than a static property of weights. As agents take on real-world responsibilities, they must transition from passive generators to self-aware systems capable of metacognition and verifiable execution.
- Recursive Agentic Reasoning and Meta$^n$: Recursive Self-Improvement through Emergent Depth highlight the power of recursive branching, while Parason: Revealing Subtask and Trial Parallelism in LLM Reasoning and Selective Regenerative Decoding: Trajectory-Level Intervention for Inference-Time Reasoning optimize the reasoning process.
- Knowing When to Ask for Help: Bayesian Self-Escalation in Hierarchical LLM Agents, IAPO: Influence-Aware Policy Optimization for Credit Assignment in Multi-Turn Service Agents, and When May an Agent Stop? Evidence-Carrying Termination for Tool-Using LLMs focus on credit assignment and verifiable termination.
- DeepRepoQA: Code Repository Question Answering with Deep Agent Exploration, How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks, ARGUS: Theory-of-Mind Guided Argument Generation with Strategy-Aware Planning and Knowledge Grounding, and LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks address the challenges of multi-hop reasoning and long-horizon planning.
- From Causal Plausibility to Causal Reliability: Evaluating LLMs as Calibrated Direct Causal-Edge Classifiers warns against treating LLMs as direct evidence, emphasizing the need for calibration.
Theme 4: Safety, Alignment, and Embodied Intelligence
As AI systems interact with the physical world and high-stakes environments, safety must evolve from post-hoc filtering to dynamic, step-level oversight.
- RePolicy: Reinforcement Learning for Safety-Policy Invocation in Agent Safeguards and StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing implement proactive safety.
- Poisoning Agentic Alpha: Adversarial Vulnerabilities Across Roles and Architectures in Multi-Agent Trading Systems and TrustShiftProbe: Characterizing, Benchmarking, and Defending Staged Trust Attacks on MCP Servers expose the vulnerabilities of agentic communication.
- Open Rubric System: Scaling Reinforcement Learning with Pairwise Adaptive Rubric, MedPriv-Bench: Benchmarking the Privacy-Utility Trade-off of Large Language Models in Medical Open-Ended Question Answering, GR-SAP: Generative Replay for Safety Alignment Preservation during Fine-Tuning, and CASTLE: A Comprehensive Benchmark for Evaluating Student-Tailored Personalized Safety in Large Language Models address the multi-dimensional nature of alignment and privacy.
- GeoWAM: Visual Geometry World Action Models for Autonomous Driving, GlanceWAM: Sparse Test-Time Imagination for World-Action Models, and NeoWorld-Pro: Programming Interactive Scenes from Monocular Images for Embodied Simulation advance the state of embodied intelligence.
Theme 5: Continual Learning & Scientific Autonomy
The future of AI lies in systems that can maintain stable representations over time and act as autonomous scientific partners capable of discovery.
- GAP-Prompt: Gated Adaptive Prompting for Efficient Continual Learning, Restoring Without Forgetting: Continual Learning Across Image Degradations, Taming foundation model with invariance-oriented pre-training for broad-spectrum EEG analysis, and Provenance Guided Incremental Learning Under Evolving Concept Definitions provide strategies for lifelong learning.
- Joint Optimization of Tool Creation and Use for Large Language Model Agents, PHMForge: Evaluating LLM Agents on Industrial Prognostics through MCP-Native, Algorithm-Grounded Tools, Self-Evolving Scientific Agent Designs Physically-Reasoned Whitebox Fluid Control, and Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models showcase the transition to scientific autonomy.
- Beyond Confidence: Test-Time Scaling for Multi-Turn Search Agents via Retrieval Grounding, A Judge Should Know What Changed: Construct Validity for LLM-as-a-Judge Evaluation, EviPathBench: Benchmarking Evidence Acquisition and Reasoning in Vision-Language Models for Whole-Slide Pathology, TRACE: An Evidence-Grounded Benchmark for Safety Evaluation of Large Reasoning Models, and ConstructCIE: A Dataset for Extracting Causal Information from Construction Accident Narratives emphasize the necessity of evidence-grounded, consequence-aware evaluation.