ArXiV ML/AI/CV papers summary
Theme 1: Physics-Informed and Structure-Aware Modeling
We are witnessing a profound shift from “black-box” statistical correlation toward models that respect the fundamental architecture of reality. By embedding physical laws—such as energy conservation, entropy production, and geometric constraints—directly into neural operators, we ensure that our models remain grounded in the laws of thermodynamics and physical dynamics.
- Physics-Informed Neural Operators: A Physics-Informed Hybrid Neural Operator for Transient Magnetization Prediction in Power Magnetics and GENERIC-FNO: Embedding Energy Conservation and Entropy Production into Fourier Neural Operators enforce conservation laws directly into operator learning.
- Symbolic Discovery: Deep Divide-and-Reduce in Symbolic Regression and Verifier-Guided Model Discovery for Physical Dynamical Systems with Pretrained Symbolic Transformers move beyond brute-force search toward theoretically grounded, verifier-guided discovery.
- Molecular and Biological Modeling: ED-DiT: Physics-Guided Diffusion Pretraining for Transferable Molecular Representations from Electron Density and TriGlue: a Biology-Inspired Generative Model for Generating Molecular Glue-Induced Ternary Complex utilize physical priors like electron density and protein-interface constraints to achieve biological plausibility.
- Geometric & 3D Grounding: Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding, Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing, and InfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View Synthesis bridge the gap between 2D pixels and 3D physical reality. NanoMorph-3D: An End-to-End Physics-Driven Unrolling Framework for Nanomaterial Reconstruction and RISE: Single Static Radar-based Indoor Scene Understanding further demonstrate the power of physics-driven reconstruction in specialized domains.
Theme 2: Agentic Reasoning, Memory, and Self-Evolution
The field is evolving from static, stateless models into autonomous agents capable of long-horizon planning, persistent memory, and self-improvement. This transition requires moving beyond simple “next-token” prediction to systems that can simulate environments and verify their own reasoning.
- Agentic Frameworks & Memory: Metas: Memory Foundation Model introduces “native memory” within the model backbone. Self-Improving Large Language Models via Progressive Experience Evolution and CrystalMem: Elastic Memory for Self-Evolving LLM Agents via Knowledge Crystallization explore how agents crystallize experiences, while Memory Reward Inflation in Self-Improving LLM Agents warns of the “Echo Gap”—a feedback loop of reinforced errors.
- Reasoning & Tool Orchestration: Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving, Learning to Coordinate Symbolic Tools: LLM Agents for Verified Sum-of-Squares Certificates, and MechGeo: Autoformalizing and Proving Euclidean Geometry in Lean 4 demonstrate agents that invoke symbolic solvers to provide mathematically checkable proofs.
- World Modeling: Quo Vadis, World Modeling?, SUV: Future Scene Understanding as Video Generation for End-to-End Driving, and Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control treat world models as interactive proxies that simulate dynamics to enable planning.
- Streaming & Continual Learning: AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?, ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?, and AiFlow: Token-Native Reactive Orchestration with Bounded Backpressure for Streaming LLM Applications address the challenges of non-isolated, continuous environments.
Theme 3: Efficient Inference and Hardware-Aware Learning
As models scale, the “plumbing” of AI—memory bandwidth, energy consumption, and hardware utilization—becomes the primary constraint. We are seeing a move toward co-designing algorithms with the underlying hardware to achieve massive throughput gains.
- KV Cache & Memory Optimization: AnchorKV: Anchor-Residual KV Cache Compression, SAKI: Score-Aware Low-Rank Key Indexing for Long-Context KV Retrieval, Output-Aware Rotation for INT2 KV-Cache Quantization, Key-Value Means: Transformers with Expandable Block-Recurrent Compressed Memory, and Keyless Attention: Value-Space Routing and Value-Only Caching for Efficient Transformers provide sophisticated methods to manage the memory bottleneck.
- Quantization & Pruning: NANQ: Noise-Floor-Aware Mixed-Precision Non-Uniform Quantization for Analog Compute-in-Memory, IPPRO: Importance-based Pruning with PRojective Offset for Magnitude-indifferent Structural Pruning, SlimVLM: Sensitivity-aware Dynamic Structured Pruning, and ActQuant: Sub-4-bit Action-Guided Quantization for Vision-Language-Action Models push the limits of model compression.
- Hardware-Algorithm Co-design: Nova: An End-to-End MLIR Compiler for Deep Learning, AgentCompile: An LLM-Guided Compiler for Direct CUDA Inference, and CN101 - A Digital Thermodynamic Computer for Generative AI represent a departure from standard CMOS architectures toward hardware-native efficiency.
Theme 4: Robustness, Safety, and Mechanistic Interpretability
To deploy AI in high-stakes domains, we must move from reactive patching to proactive, verifiable safety. This involves mapping the internal geometry of models to understand how they represent truth and uncertainty.
- Mechanistic Interpretability: Language Models Encode the Contextual Truth of Propositions, States Hidden in Hidden States: Implicit Discrete State Representations Emerge in LLMs’ Hidden States, and The Transformer as a Polar State Estimator reveal the internal symbolic and geometric structures that govern model behavior.
- Safety & Alignment: Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety, AI Security Leaderboard: Methodology, Results and Minimal Standard, and Risky Business: Measuring The Faithfulness-Safety Tension highlight the critical trade-offs between model faithfulness and safety.
- Unlearning & Auditing: Trajectory-Guided Forget-Recover Network for Continual LLM Unlearning, One-Point Contraction: Erasing Representational Separability toward Irreversible Deep Forgetting, and UHP Detection: LVLMs have their Unique Hallucination Pattern in the Consistency Space provide tools for verifiable data removal and hallucination detection.
Theme 5: Advanced Optimization and Human-Centric Paradigms
The final frontier involves refining the fundamental optimization algorithms and ensuring that AI systems align with human values and societal needs.
- Spectral & Geometric Optimization: Joint Affine Spectral Shaping: Coupling Weight and Bias Updates Beyond Weight-Only Muon, Muon Meets Mamba: Spectral Optimization for State Space Models, and Information-Geometric Forward Policy Training in GFlowNets explore how the geometry of parameter space guides stable training.
- Conformal Prediction: Conformal Anomaly Detection in Python: Moving Beyond Heuristic Thresholds with nonconform and Online Shift Detection and Conformal Adaptation for Deployed Safety Classifiers provide rigorous, uncertainty-aware decision-making tools.
- Societal Impact: Me and My Bot: What Users Talk About in AI Companion Communities on Reddit, CompanionBench: A Theory-Anchored, Real-World-Grounded Benchmark for AI Emotional Companionship, and The Epistemic Politics of AI Anthropomorphism remind us that AI development is as much a sociological endeavor as a technical one.