Research 10.0 score cs.AI
Adam Karvonen, Euan Ong, et al.
Flagged for: agent, agentic, chain of thought, interpretability
agentagenticchain of thoughtinterpretabilityinterpretllm
Many areas of AI research, such as language model interpretability and chain of thought faithfulness, seek to explain model behaviors. But what constitutes a "good" explanation? In this work, we evalu...
Robotics 10.0 score cs.RO
Kangning Yin, Kaige Liu, et al.
Flagged for: agi, agent, multi-agent, autonomous
agiagentmulti-agentautonomoushumanoidrobot
Achieving human-level competitive intelligence and physical agility in humanoid robots remains a profound challenge, particularly in contact-rich and highly dynamic tasks such as boxing. While Multi-A...
Research 10.0 score cs.CL
Mason Smetana, Trevor Neece, et al.
Flagged for: agent, agentic, planning, ai safety
agentagenticplanningai safetyllmlanguage model
Job Safety Analysis (JSA) and pre-task planning can benefit from prior incident records, yet historical accident data is often stored as unstructured narratives that are difficult to consult at the po...
Robotics 9.5 score cs.RO
Jiawei Liu, Jiacheng Guo, et al.
Flagged for: agent, reasoning, embodied, robot
agentreasoningembodiedrobotllmlanguage model
Large Language Models (LLMs) have demonstrated capabilities in in-context learning, task decomposition, step-by-step reasoning, and code generation, driving their gradual evolution from text generatio...
Research 9.5 score cs.AI
Hongyue Yu, Kefan Li, et al.
Flagged for: agi, agent, reasoning, llm
agiagentreasoningllmlanguage modelrag
Large Language Models (LLMs) have achieved remarkable progress in code generation, yet ensuring correctness in complex, repository-level tasks remains challenging. Existing approaches often use genera...
Research 9.0 score cs.AI
Reza Fayyazi, Michael Zuzak, et al.
Flagged for: agent, agentic, autonomous, llm
agentagenticautonomousllmlanguage modelrag
Large Language Models (LLMs) are increasingly being deployed in cybersecurity operations to assist cybersecurity analysts with rapid decision-making against emerging threats. However, there is a main ...
Research 8.5 score cs.CL
Hao Zhang, Longrong Yang, et al.
Flagged for: agent, multi-agent, agentic, scaling
agentmulti-agentagenticscalingretrievalrag
Multi-modal retrieval-augmented generation (RAG) is a key technique for visually rich long document understanding. Existing multi-modal RAG methods are progressively advancing toward multi-agent syste...
Research 8.5 score cs.CL
Ruiyao Xu, Tiankai Yang, et al.
Flagged for: agi, agent, agentic, llm
agiagentagenticllmrag
As agentic tasks grow in complexity, LLM agents increasingly rely on experiential memory to reuse procedural knowledge across tasks. Effective memory design must jointly address what to store, how mem...
Robotics 8.5 score cs.RO
Vineet Bhat, Siyi Chen, et al.
Flagged for: agent, agentic, planning, reasoning
agentagenticplanningreasoningrobotlanguage model
Vision-language Models (VLMs) excel at 2D grounding, spatial reasoning and agentic tool-based planning in static scenes. However, consider asking a home robot "Is my medication still in the cabinet?" ...