ArXiV ML/AI/CV papers summary
This collection of research represents a vibrant, multi-disciplinary frontier in machine learning. As we push the boundaries of artificial intelligence, we are moving beyond simple pattern matching toward systems that reason, verify, and interact with the physical and logical world. By treating models as objects of scientific study—applying the rigor of geometry, physics, and information theory—we are transitioning from “black-box” engineering to a principled understanding of intelligence.
Theme 1: Mechanistic Interpretability and Geometric Foundations
A central challenge in modern AI is moving from opaque performance to a rigorous understanding of why models behave as they do. Researchers are increasingly using tools from geometry and physics to map internal logic.
- Internal Logic & Truth: Latent Fact-Checking: Detecting Misinformation through Activation Engineering reveals that truthfulness is a geometric property within latent space. Similarly, Mathematical Principles and Experimental Discoveries of the Emergence of Symbolic Patterns in Artificial Neural Networks suggests that explainability is a natural law of deep learning, as networks converge toward sparse, symbolic interactions.
- Reverse Engineering Structures: Cryptanalytic Extraction of Isolated Bias-Free GLU Feed-Forward Blocks by Antipodal Separation and Hidden Gauge Controls Feature Specialization in ReLU Networks provide methods to deconstruct internal network blocks and neuron specialization.
- Geometric Learning: The Geometric Mechanics of Contrastive Representation Learning: Alignment Potentials, Entropic Dispersion, and Cross-modal Divergence and Wasserstein Mahalanobis Distances for Recovering Latent Geometry shift our perspective toward the energy landscapes and manifolds that define how models learn.
Theme 2: Agentic Reasoning, Planning, and Memory
We are witnessing a shift from models that “answer” to agents that “act.” This requires maintaining state, planning over long horizons, and correcting errors in real-time.
- Memory Management: To prevent “memory pollution” and stale evidence, TEPA: Revoking Stale Memories for Conflict-Robust Language Agents and Blast Radius introduce mechanisms to revoke or bury dead context. LiveMem: Maintaining Memory State Continuity in Long-Running LLM Inference and Toward Reliable Context Compression for Long-Horizon Agents: An Empirical Study of Execution Instability address the fundamental bottleneck of maintaining focus over long interaction streams.
- Planning & Coupling: PMCoder: Coupling Planning with Episodic Memory in LLM Agents for Software Issue Resolution and Recursive Synthesis for Long-Horizon Terminal Tasks demonstrate that coupling planning with memory and recursive verification is essential for complex, multi-step tasks. Sharding Prevents LLM Oversight Failures and Adversarial Exploitation further improves reliability by breaking tasks into manageable shards.
- Affective & Social Cognition: PsychoAgent: An Affect-Sensitive Cognitive Architecture for Conflict-Aware Memory in LLM Agents and Social World Models explore how agents can prioritize emotionally salient traces and reason about the mental states of others.
Theme 3: Agentic RL and Credit Assignment
As agents become autonomous, determining which specific actions contribute to success—the “credit assignment” problem—has become a primary focus.
- Fine-Grained Credit: DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training and FACTOR: Credit-Conserving Action-to-Token Allocation for Multi-Turn Agent Reinforcement Learning construct credit units from structural data (like code diffs) to prevent bias.
- On-Policy Distillation: MemOPD: On-Policy Distillation through Memory State Alignment for Long-Horizon Agents and DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models provide methods to maintain teacher supervision validity even when memory is compressed or reasoning diverges over time.
Theme 4: Scientific and Domain-Specific Foundation Models
AI is increasingly simulating the physical and clinical world, where accuracy and provenance are non-negotiable.
- Physics & Simulation: Fixed and Adaptive Topological DeepONets: Functional Measurements on Hausdorff Locally Convex Spaces, Fluid-DiT: Graph-Free Diffusion Transformers for Fluid Flow Simulations Learning, and Unsupervised Adaptation of PDE Foundation Models push the boundaries of simulating physical systems. SoRoMoX: Fast, Differentiable, and Parallelizable Soft Robot Models and AutoMOOSE: An Agentic AI for Autonomous Phase-Field Simulation automate complex physical workflows.
- Material & Biological Discovery: ED-CSP: Crystal Structure Prediction from Electron Diffraction, DynaCrys: Crystal Generation with Dynamic Space-Group Diffusion, and GeneGeoFlow apply generative models to chemistry and biology. Genotypic Triggers warns of security risks in these generative pipelines.
- Clinical & Financial Provenance: ResidencyRL: Reinforcement Learning in Simulated Clinical Environments, FinRank: An Evidence-Grounded Benchmark for Financial Question Answering and Retrieval over SEC Filings, and READ: Reliable Embedding-free Agentic Document-search emphasize audit-trail-based operations over opaque retrieval. Adversarial Causal Intervention Falsification ensures models encode correct causal structures.
Theme 5: Efficiency, Quantization, and Deployment
As models grow, the engineering constraints of energy, latency, and memory have become first-class research objectives.
- Inference Optimization: Quantization Damage Is Multiplicative, Not Additive and CubicQuant: Parametric Non-Uniform Codebooks for High-Throughput LLM Inference with 1-8-Bit Weights provide deep insights into bit-width reduction. Autonomy-of-Heads: Data-Free Sparse Attention from Frozen Query-Key Geometry and CoBa: Cost-Effective Test-Time Scaling via Compute-Balanced Routing optimize the inference budget.
- Efficiency Metrics: Intelligence per Watt: Measuring Intelligence Efficiency of Local AI introduces a critical metric for the field, while DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts demonstrates kernel-level optimizations for training.
- Specialized Quantization: PTQ4SNN: Membrane-Aware Post-Training Quantization for Spiking Neural Networks and LoRAScan address specific challenges in spiking networks and LoRA-based security.
Theme 6: Vision-Language Models and Medical Imaging
The expansion of VLMs requires intelligent compression, while medical AI requires hierarchical structural priors.
- VLM Efficiency: Direct Visual Grounding by Directing Attention of Visual Tokens, Decoupled Similarity for Task-Aware Token Pruning in Large Vision-Language Models, and VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression move away from brute-force processing. A Picture is Worth a Thousand Tokens demonstrates that encoding time-series as 2D plots reduces token counts.
- Medical Hierarchies: Learning Coarse-to-Fine Osteoarthritis Representations under Noisy Hierarchical Labels, Thinking in Scales: Accelerating Gigapixel Pathology Image Analysis via Adaptive Continuous Reasoning, and Curia-MAE: Multi-Modal Multi-Anatomy MAE Pre-Training for 3D Medical Image Segmentation bake clinical realities into model architectures.
- Adaptive Perception: SynthRender and I-AsSET, An active-learning framework for real-time depth perception from monocular vision streams, and YOLOv14: Adaptive Real-Time Object Detection for Diverse Imaging Conditions focus on robustness in the “wild.”
Theme 7: Safety, Auditing, and Alignment
As systems are deployed in high-stakes environments, ensuring robustness and alignment is paramount.
- Safety & Auditing: Bypassing Krum: Selection-Aware Backdoor Attacks in Federated Learning, Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits, A2E: An End-to-End Agent Auditing Engine, and HarnessSafe highlight the ongoing arms race in AI safety. NiyamAI offers a path toward provable safety using zero-knowledge proofs.
- Alignment & Fairness: CASA: Classification Augmented with Safety Attention for Robust Multimodal Alignment and Let’s Unlearn Stereotypes Before Decision-Making provide methods for internal bias mitigation. People Are Not Just Their Countries and Do AI Personas Grow? explore the socio-demographic and psychological dimensions of alignment.
Theme 8: Efficient Generative Modeling
The frontier of generative modeling is making high-fidelity diffusion models “one-step” or “real-time.”
- Distillation & Trajectory: Teacher-Feature Drifting: One-Step Diffusion Distillation with Pretrained Diffusion Representations, CloudDiffusion: Diffusion-Based Scene Completion in the Point Cloud Domain, and Energy-Guided Flow Matching prove that high-fidelity generation can be achieved through smarter scheduling and trajectory evolution rather than just increased compute.