07. IA agentica
updated
Paper
• 2606.01533
• Published • 8
OpenSkill: Open-World Self-Evolution for LLM Agents
Paper
• 2606.06741
• Published • 29
Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills
Paper
• 2606.07412
• Published • 12
Bayesian-Agent: Posterior-Guided Skill Evolution for LLM Agent Harnesses
Paper
• 2606.08348
• Published • 16
SlimSearcher: Training Efficiency-Aware Web Agents via Adaptive Reward Gating
Paper
• 2606.07074
• Published • 12
Lean4Agent: Formal Modeling and Verification for Agent Workflow and Trajectory
Paper
• 2606.06523
• Published • 7
Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution
Paper
• 2606.10917
• Published • 76
Retrospective Harness Optimization: Improving LLM Agents via Self-Preference over Trajectory Rollouts
Paper
• 2606.05922
• Published • 71
Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models
Paper
• 2606.11025
• Published • 41
Rethinking the Divergence Regularization in LLM RL
Paper
• 2606.09821
• Published • 34
Online Skill Learning for Web Agents via State-Grounded Dynamic Retrieval
Paper
• 2606.04391
• Published • 11
What Should Agents Say? Action-state Communication for Efficient Multi-Agent Systems
Paper
• 2606.05304
• Published • 5
Struct-Searcher: Agentic Structural Thinking Advances Multimodal Deep Information Seeking
Paper
• 2606.07689
• Published • 6
Paper
• 2606.10650
• Published • 10
Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application
Paper
• 2606.12191
• Published • 72
Toward Generalist Autonomous Research via Hypothesis-Tree Refinement
Paper
• 2606.11926
• Published • 130
EvoBrowseComp: Benchmarking Search Agents on Evolving Knowledge
Paper
• 2606.13120
• Published • 4
WebChallenger: A Reliable and Efficient Generalist Web Agent
Paper
• 2606.10423
• Published • 2
Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling
Paper
• 2606.12370
• Published • 21
Large Language Models Are Overconfident in Their Own Responses
Paper
• 2606.03437
• Published • 3
HarnessBridge: Learnable Bidirectional Controller for LLM Agent Harness
Paper
• 2606.12882
• Published • 15
The Arbiter Agent: Continually Monitoring Multi-Agent Conversations to Detect Emergent Misalignment
Paper
• 2606.10747
• Published • 13
LLM Agents Can See Code Repositories
Paper
• 2606.14061
• Published • 20
Skip a Layer or Loop It? Learning Program-of-Layers in LLMs
Paper
• 2606.06574
• Published • 25
HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry
Paper
• 2606.14249
• Published • 50
APPO: Agentic Procedural Policy Optimization
Paper
• 2606.12384
• Published • 78
CODA-BENCH: Can Code Agents Handle Data-Intensive Tasks?
Paper
• 2606.15300
• Published • 13
FastContext: Training Efficient Repository Explorer for Coding Agents
Paper
• 2606.14066
• Published • 96
PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems
Paper
• 2606.22388
• Published • 96
DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams
Paper
• 2606.21337
• Published • 75
CLI-Universe: Towards Verifiable Task Synthesis Engine for Terminal Agents
Paper
• 2606.22883
• Published • 37
Learning from Your Own Mistakes: Constructing Learnable Micro-Reflective Trajectories for Self-Distillation
Paper
• 2606.18844
• Published • 20
Self-Compacting Language Model Agents
Paper
• 2606.23525
• Published • 19
Causal Discovery in the Era of Agents
Paper
• 2606.23608
• Published • 8
FastMix: Fast Data Mixture Optimization via Gradient Descent
Paper
• 2606.14971
• Published • 3
DanceOPD: On-Policy Generative Field Distillation
Paper
• 2606.27377
• Published • 83
OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning
Paper
• 2606.26790
• Published • 57
The Verification Horizon: No Silver Bullet for Coding Agent Rewards
Paper
• 2606.26300
• Published • 53
PhysiFormer: Learning to Simulate Mechanics in World Space
Paper
• 2606.27364
• Published • 11
When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models
Paper
• 2606.27288
• Published • 4
Are We Ready For An Agent-Native Memory System?
Paper
• 2606.24775
• Published • 134
Plans Don't Persist: Why Context Management Is Load Bearing for LLM Agents
Paper
• 2606.22953
• Published • 1
Forecasting Future Behavior as a Learning Task
Paper
• 2606.11445
• Published • 1
Qwen-AgentWorld: Language World Models for General Agents
Paper
• 2606.24597
• Published • 160
Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning
Paper
• 2606.24428
• Published • 52
OpenThoughts-Agent: Data Recipes for Agentic Models
Paper
• 2606.24855
• Published • 48
Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMs
Paper
• 2606.27378
• Published • 60
Towards Automating Scientific Review with Google's Paper Assistant Tool
Paper
• 2606.28277
• Published • 11
Cluster, Route, Escalate: Cascaded Framework for Cost-Aware LLM Serving
Paper
• 2606.27457
• Published • 5
MemSyco-Bench: Benchmarking Sycophancy in Agent Memory
Paper
• 2607.01071
• Published • 31
Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning
Paper
• 2607.00461
• Published • 29
CausalMix: Data Mixture as Causal Inference for Language Model Training
Paper
• 2607.01104
• Published • 21
The State-Prediction Separation Hypothesis
Paper
• 2607.01218
• Published • 12
AutoTrainess: Teaching Language Models to Improve Language Models Autonomously
Paper
• 2606.31551
• Published • 24
Valdi: Value Diffusion World Models
Paper
• 2607.00917
• Published • 15
Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?
Paper
• 2607.01211
• Published • 14
Autonomous Scientific Discovery via Iterative Meta-Reflection
Paper
• 2607.01131
• Published • 8
When LLMs Read Tables Carelessly: Measuring and Reducing Data Referencing Errors
Paper
• 2606.32029
• Published • 14
Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks
Paper
• 2606.29082
• Published • 43
Managing Procedural Memory in LLM Agents: Control, Adaptation, and Evaluation
Paper
• 2606.23127
• Published • 26
Little Brains, Big Feats: Exploring Compact Language Models
Paper
• 2606.30062
• Published • 16
Hierarchical Experimentalist Agents
Paper
• 2606.29315
• Published • 5
SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Use
Paper
• 2607.01874
• Published • 23
PACE: A Proxy for Agentic Capability Evaluation
Paper
• 2607.02032
• Published • 20
When Search Agents Should Ask: DiscoBench for Clarification-Aware Deep Search
Paper
• 2606.27669
• Published • 16
Paper
• 2607.27201
• Published • 109
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
Paper
• 2608.05987
• Published • 101
ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment
Paper
• 2608.05102
• Published • 68
Weak-to-Strong On-Policy Distillation
Paper
• 2607.26246
• Published • 58
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes
Paper
• 2608.05000
• Published • 62
Progressive Agent Skill Generation via Reinforcement Learning
Paper
• 2608.01678
• Published • 60
HarnessOpt-Bench: Evaluating LLMs at Harness Optimization
Paper
• 2608.06301
• Published • 35
WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity
Paper
• 2608.02603
• Published • 34
OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents
Paper
• 2608.05013
• Published • 37
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks
Paper
• 2608.03764
• Published • 27
Continual Learning in Transition
Paper
• 2608.06216
• Published • 26
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?
Paper
• 2608.00155
• Published • 25
CADENA: Stepwise CAD Reverse Engineering
Paper
• 2608.00799
• Published • 41
Kimi K3: Open Frontier Intelligence
Paper
• 2607.24653
• Published • 514
Program-as-Weights: A Programming Paradigm for Fuzzy Functions
Paper
• 2607.02512
• Published • 309
Metis: Memory Foundation Model
Paper
• 2607.26760
• Published • 273
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable
Paper
• 2607.13285
• Published • 235
AREX: Towards a Recursively Self-Improving Agent for Deep Research
Paper
• 2607.21461
• Published • 154
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines
Paper
• 2607.16617
• Published • 144
Weak-to-Strong Generalization via Direct On-Policy Distillation
Paper
• 2607.05394
• Published • 150
EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World
Paper
• 2607.17250
• Published • 93
Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation
Paper
• 2607.05382
• Published • 88
ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog
Paper
• 2607.04438
• Published • 64
Skaling: Chinchilla's Exponents Meet Kaplan's Coupling
Paper
• 2608.07222
• Published • 11
Round-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors
Paper
• 2608.00675
• Published • 10
The Optimizer Is the Agent: Reasoning-Driven Search across Prompts, Programs, and ML Workflows
Paper
• 2608.06714
• Published • 10
Can MLLMs Decode the Creative Leap? Introducing C4 for Cross-Concept Understanding
Paper
• 2608.06501
• Published • 3
CEAA: A Cognitive Embodied Agents Architecture for Interactive Computing Systems
Paper
• 2608.09848
• Published • 6
Don't Scroll Back: Missing-Evidence Memory for Streaming Dialogue Summarization
Paper
• 2608.09043
• Published • 8
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning
Paper
• 2608.09888
• Published • 770
Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory
Paper
• 2608.07169
• Published • 50
Motif 3: Technical Report
Paper
• 2608.09119
• Published • 46
Scaling Inherently Interpretable Language Models
Paper
• 2608.07594
• Published • 21
Stealing Reasoning Traces from Proprietary LLM APIs
Paper
• 2608.09867
• Published • 117
Evo-Bench: Can Language Models Improve Agent Harness?
Paper
• 2608.09096
• Published • 18
ComBodied Agents: a New Paradigm of Human-Centric Agentic AI
Paper
• 2608.10915
• Published • 195
Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design
Paper
• 2608.10299
• Published • 135
Mendel Gödel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution
Paper
• 2608.07645
• Published • 27
VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?
Paper
• 2608.10875
• Published • 17
Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness
Paper
• 2608.09900
• Published • 13
SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information
Paper
• 2608.10692
• Published • 12
Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents
Paper
• 2608.08389
• Published • 12
DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation
Paper
• 2608.10636
• Published • 11
DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?
Paper
• 2608.10366
• Published • 10
Gaze Target Estimation Anywhere with Concepts
Paper
• 2608.11367
• Published • 5
SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries
Paper
• 2608.05604
• Published • 80
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives
Paper
• 2608.08160
• Published • 29
Self-Evolving Embodied Agents via Skill-Harness Evolution
Paper
• 2608.11350
• Published • 13
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses
Paper
• 2608.12307
• Published • 114
DarwinX: Evolving Agent Harnesses Through Natural Selection
Paper
• 2608.07545
• Published • 113
AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design
Paper
• 2608.13560
• Published • 55
Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence
Paper
• 2608.12743
• Published • 44
LycheeMemory V2: Efficient Long-Term Memory for LLM Agents via Semantic Segment-Level Consolidation
Paper
• 2608.12990
• Published • 14
Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning
Paper
• 2607.29211
• Published • 14
Marionette: Predicting World States, Rendering Geometry, Painting Appearance
Paper
• 2608.14530
• Published • 33
Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence
Paper
• 2608.11341
• Published • 66
Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning
Paper
• 2608.14290
• Published • 33
Second Thought: Reasoning in Parallel as LLM Agents Act and Observe
Paper
• 2608.13667
• Published • 16
Multimodal Model Diffing for Feature Discovery and Control
Paper
• 2608.09928
• Published • 10
Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems
Paper
• 2608.03744
• Published • 6
A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images
Paper
• 2608.14075
• Published • 4
UniProbe: A Learnable Token-Level Hallucination Detector for Large VLMs using Multi-Structural Internal Representations
Paper
• 2608.10835
• Published • 4
StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling
Paper
• 2608.15089
• Published • 445
Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search
Paper
• 2608.15669
• Published • 62
Demystifying Agent Skills: Why They Work-Until They Don't
Paper
• 2608.14036
• Published • 169
Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements
Paper
• 2608.17310
• Published • 107
Agent Lightning v1.0: Towards Harnessed Agentic RL
Paper
• 2608.17528
• Published • 32
SkillForge: Self-Distilling Agents for Project-Specific Issue Resolution
Paper
• 2608.18933
• Published • 12
HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety
Paper
• 2608.17597
• Published • 10
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
Paper
• 2608.17800
• Published • 9
From Sequence to Structure: Relational Uncertainty Propagation for LLM Agents
Paper
• 2608.16002
• Published • 7
The More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning
Paper
• 2608.14229
• Published • 17
LLMs Get Smarter from Targeted Synthetic Multilingual Data
Paper
• 2608.15964
• Published • 5
The Embedder's Dilemma: LLMs Are Better, but at What Cost?
Paper
• 2608.12875
• Published • 15
EnvHarness: Awakening Static Worlds for Agent Learning
Paper
• 2608.19880
• Published • 274
TinyCast: Probabilistic Zero-Shot Forecasting with Computed Periodicity
Paper
• 2608.15767
• Published • 9
FlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills
Paper
• 2607.21596
• Published • 21
AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces
Paper
• 2608.23041
• Published • 64
Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses
Paper
• 2608.24876
• Published • 28
CAFE: Self-Improving Search Agents Need Co-Evolving Feedback
Paper
• 2608.24794
• Published • 6
Automata from Agent Traces: Failure and Next-Step Prediction
Paper
• 2608.23670
• Published • 5
Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment
Paper
• 2608.23691
• Published • 3
JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution
Paper
• 2608.25593
• Published • 115
FrontierChallenge: Evaluating Scientific Workflow Completion
Paper
• 2608.24979
• Published • 146
Agent-G^2: Gaussian Guidance for Agentic Reinforcement Learning
Paper
• 2608.23318
• Published • 29
Code World Model: Coding Agent as World Brain
Paper
• 2608.25927
• Published • 35
Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios
Paper
• 2608.25529
• Published • 16
A Modular Agent for Reliable and Auditable Spatial Relation Verification in CT Scans
Paper
• 2608.21140
• Published • 2
Real-TurnTurk: A Multimodal Turkish Corpus for Turn-Taking Prediction
Paper
• 2608.22071
• Published
PAWBench: How Far Are We from Probabilistically Aligned World Modeling?
Paper
• 2608.27345
• Published • 143
Training Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical Report
Paper
• 2608.15763
• Published • 50
CaRGo-T: Causal Reasoning Graph-of-Thought improves Multimodal Humor Comprehension
Paper
• 2608.23172
• Published • 9