https://www.academia.edu/3071-0286/2/1/10.20935/AcadAI8229
Introduction: Meta-reasoning, the ability of an autonomous agent to monitor and regulate its own reasoning processes, has emerged as a critical component for achieving adaptive and trustworthy artificial intelligence. However, limited empirical evidence quantifies how meta-reasoning impacts agent performance across diverse contexts and model architectures. This study presents a comprehensive empirical analysis of meta-reasoning capabilities in large language model (LLM)-based autonomous agents using two established benchmarks: GAIA (General AI Assistants) and AgentBench.
Materials and methods: We evaluate five state-of-the-art (SOTA) models spanning proprietary and open-source architectures: GPT-4o, Claude 3.5 Sonnet, Gemini 2.0 Flash, Llama 3.3 70B, and Qwen 2.5 72B through controlled experiments using a Meta-reasoning control loop involving 1165 tasks across reasoning, web browsing, and tool-use domains.
Results: We demonstrate that agents equipped with meta-reasoning frameworks exhibit statistically significant improvements. Results show a 31.2% (p < 0.001) increase in task completion rates, 24.7% improvement in decision quality, and 18.9% reduction in API token consumption compared to baseline agents. Notably, meta-reasoning benefits generalize across all five model architectures, with the largest improvements observed in multi-step reasoning tasks.
Conclusions: These findings establish meta-reasoning as a model-agnostic enhancement that yields the largest benefits for complex, failure-prone agentic workflows.
No hay comentarios:
Publicar un comentario