<?xml version="1.0" ?>
<rss xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title>AI Paper Daily</title>
    <description>Daily AI Paper Discovery · Agent / RAG / Knowledge Graph</description>
    <link>https://alloevil.github.io/AI-Paper-Daily</link>
    <language>zh-CN</language>
    <lastBuildDate>Sun, 16 Aug 2026 04:41:59 +0000</lastBuildDate>
    <atom:link href="https://alloevil.github.io/AI-Paper-Daily/feed.xml" rel="self" type="application/rss+xml"/>
    <item>
      <title>📄 Beyond Handcrafted Security: Towards Self-Evolving Defense for LLM Agents (2026-08-16)</title>
      <description>The expanding operational capabilities of large language model (LLM) agents introduce sophisticated security threats. Runtime defenses have emerged as an effective approach to mitigating these risks b...</description>
      <link>http://arxiv.org/abs/2608.12977v1</link>
      <guid>http://arxiv.org/abs/2608.12977v1#2026-08-16</guid>
      <pubDate>Sun, 16 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 Discovering Efficient and Explainable Communication Topologies for LLM-based Multi-Agent Systems via Causal Inference (2026-08-16)</title>
      <description>The performance of large language model (LLM)-based multi-agent systems (MAS) largely depends on effective communication topologies. Existing topology generation methods, however, typically learn comm...</description>
      <link>http://arxiv.org/abs/2608.12921v1</link>
      <guid>http://arxiv.org/abs/2608.12921v1#2026-08-16</guid>
      <pubDate>Sun, 16 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 When Your Agent Opens the Chat App: Agent-Controlled Search over Raw Chat Logs Rivals Structured Memory (2026-08-16)</title>
      <description>Agent-memory systems increasingly buy retrieval quality with structure, transforming raw conversation histories into summaries, embeddings, trees, or knowledge graphs before any question is asked. We ...</description>
      <link>http://arxiv.org/abs/2608.12888v1</link>
      <guid>http://arxiv.org/abs/2608.12888v1#2026-08-16</guid>
      <pubDate>Sun, 16 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents (2026-08-16)</title>
      <description>Self-improving LLM agents convert successful trajectories into persistent cross-task state. An unsafe success can thereby become reusable policy after its triggering input disappears. Skill evolution ...</description>
      <link>http://arxiv.org/abs/2608.12851v1</link>
      <guid>http://arxiv.org/abs/2608.12851v1#2026-08-16</guid>
      <pubDate>Sun, 16 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 AaLLM: An End-to-End Analog Circuit Design Framework from Topology Generation to Sizing Using Large Language Models (2026-08-15)</title>
      <description>Analog circuit design is a time-consuming, iterative process in a nonlinear and high-dimensional design space that relies heavily on expert intuition. Among recent developments, LLMs have introduced a...</description>
      <link>http://arxiv.org/abs/2608.13472v1</link>
      <guid>http://arxiv.org/abs/2608.13472v1#2026-08-15</guid>
      <pubDate>Sat, 15 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 StateBridge: Training-free Hidden-state Alignment for Latent Communication in LLM Multi-Agent Systems (2026-08-15)</title>
      <description>Large language model based multi-agent systems usually communicate in text, i.e., using discrete tokens. However, text introduces a discrete bottleneck. Converting the sender's continuous hidden state...</description>
      <link>http://arxiv.org/abs/2608.13317v1</link>
      <guid>http://arxiv.org/abs/2608.13317v1#2026-08-15</guid>
      <pubDate>Sat, 15 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 Novel Knowledge-Guided Generative Methods for Synthetic Transcriptomic Data (2026-08-15)</title>
      <description>As biomedical research increasingly relies on data-intensive tools, the quality and utility of datasets are critical. Challenges such as imbalances, biases, and ethical or legal constraints often limi...</description>
      <link>http://arxiv.org/abs/2608.13256v1</link>
      <guid>http://arxiv.org/abs/2608.13256v1#2026-08-15</guid>
      <pubDate>Sat, 15 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents (2026-08-15)</title>
      <description>Reinforcement learning with verifiable rewards (RLVR) offers a verifier-bounded performance ceiling for training multi-turn tool-use agents, yet its trajectory-level credit assignment conflates hetero...</description>
      <link>http://arxiv.org/abs/2608.13179v1</link>
      <guid>http://arxiv.org/abs/2608.13179v1#2026-08-15</guid>
      <pubDate>Sat, 15 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 SkillShapley: Boundary-Adaptive Shapley Valuation for Skill Step Attribution in LLM Agents (2026-08-15)</title>
      <description>Agent skills are crucial external instructions that enable language agents to execute long procedural tasks such as coding or document processing. Existing agent skills are primarily created through h...</description>
      <link>http://arxiv.org/abs/2608.13173v1</link>
      <guid>http://arxiv.org/abs/2608.13173v1#2026-08-15</guid>
      <pubDate>Sat, 15 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 Operationalizing Cyber Threat Intelligence with GraphRAG (2026-08-15)</title>
      <description>When a security researcher publishes a report on a cyberattack, detection engineers are supposed to turn it into working detection rules. In practice, most automated attempts at this only extract the ...</description>
      <link>http://arxiv.org/abs/2608.13050v1</link>
      <guid>http://arxiv.org/abs/2608.13050v1#2026-08-15</guid>
      <pubDate>Sat, 15 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 BoardroomAI: Dependency-Aware Human-Steerable Multi-Agent Deliberation through Evolving Decision Graphs (2026-08-15)</title>
      <description>Organizational decisions are co-created while evidence, constraints, and human priorities continue to evolve. In conventional transcript-based multi-agent systems, humans typically provide an initial ...</description>
      <link>http://arxiv.org/abs/2608.13046v1</link>
      <guid>http://arxiv.org/abs/2608.13046v1#2026-08-15</guid>
      <pubDate>Sat, 15 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 Static analysis-guided agentic AI translation enables Rust as a full stack bioinformatics language (2026-08-15)</title>
      <description>The field of bioinformatics struggles with legacy code - old code that is commonly used but may no longer have a maintainer, or may be written in an now-unfamiliar language (e.g. Perl, Fortran). This ...</description>
      <link>http://arxiv.org/abs/2608.13029v1</link>
      <guid>http://arxiv.org/abs/2608.13029v1#2026-08-15</guid>
      <pubDate>Sat, 15 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 HybridRAG-BN: A Retrieval-Augmented Framework with Fine-Tuned Verification for Bangla KBQA (2026-08-15)</title>
      <description>Knowledge-base question answering (KBQA) systems rely on effective retrieval and reasoning mechanisms to generate accurate answers from external knowledge sources. However, developing reliable KBQA sy...</description>
      <link>http://arxiv.org/abs/2608.13004v1</link>
      <guid>http://arxiv.org/abs/2608.13004v1#2026-08-15</guid>
      <pubDate>Sat, 15 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 LycheeMemory V2: Efficient Long-Term Memory for LLM Agents via Semantic Segment-Level Consolidation (2026-08-15)</title>
      <description>Long-horizon LLM agents must preserve information from past interactions to support future tasks. Existing memory systems typically rely on eager consolidation, invoking LLMs after each interaction to...</description>
      <link>http://arxiv.org/abs/2608.12990v1</link>
      <guid>http://arxiv.org/abs/2608.12990v1#2026-08-15</guid>
      <pubDate>Sat, 15 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation (2026-08-14)</title>
      <description>We present DreamX-Phi 1.0, an action-conditioned video world model for robotic manipulation that, given an observed frame, a language instruction, and a prescribed action sequence comprising end-effec...</description>
      <link>https://arxiv.org/abs/2608.13489</link>
      <guid>https://arxiv.org/abs/2608.13489#2026-08-14</guid>
      <pubDate>Fri, 14 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 Intern-S2-Preview: Scientific Agentic Foundation Model (2026-08-14)</title>
      <description>Scientific discovery increasingly requires AI systems that can reason over scientific evidence of heterogeneous modalities, interact with scientific tools and environments, and sustain progress across...</description>
      <link>https://arxiv.org/abs/2608.13505</link>
      <guid>https://arxiv.org/abs/2608.13505#2026-08-14</guid>
      <pubDate>Fri, 14 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence (2026-08-14)</title>
      <description>Spatial intelligence is becoming a foundation for embodied agents, robotic planning, and multimodal assistants. To improve the spatial reasoning ability of VLM agents, existing work has mainly followe...</description>
      <link>https://arxiv.org/abs/2608.12743</link>
      <guid>https://arxiv.org/abs/2608.12743#2026-08-14</guid>
      <pubDate>Fri, 14 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design (2026-08-14)</title>
      <description>Transforming multimodal sources into condensed and structured media outputs can be fundamentally conceptualized as a long-horizon agentic process centered on a model-harness system. While an ideal har...</description>
      <link>https://arxiv.org/abs/2608.13560</link>
      <guid>https://arxiv.org/abs/2608.13560#2026-08-14</guid>
      <pubDate>Fri, 14 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives (2026-08-14)</title>
      <description>Video world models simulate future states conditioned on current observations and user actions. Recent systems have demonstrated impressive video consistency and action controllability over long seque...</description>
      <link>https://arxiv.org/abs/2608.13552</link>
      <guid>https://arxiv.org/abs/2608.13552#2026-08-14</guid>
      <pubDate>Fri, 14 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence 📦代码 (2026-08-14)</title>
      <description>Compositional reliability bounds for multi-agent systems multiply component reliabilities, a step licensed by a conditional-independence assumption that is routinely stated and rarely tested. We test ...</description>
      <link>http://arxiv.org/abs/2608.12895v1</link>
      <guid>http://arxiv.org/abs/2608.12895v1#2026-08-14</guid>
      <pubDate>Fri, 14 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 Alaya-EVOKE: From Linear-Scaling Supervision to Endless World (2026-08-14)</title>
      <description>Interactive world models must support persistent memory, responsive interaction, and long-horizon generation, yet these requirements place conflicting demands on the model. Maintaining history in the ...</description>
      <link>https://arxiv.org/abs/2608.13546</link>
      <guid>https://arxiv.org/abs/2608.13546#2026-08-14</guid>
      <pubDate>Fri, 14 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models (2026-08-14)</title>
      <description>Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive and difficult to scale. Meanwhile, abundant egocentric human manipulation videos pr...</description>
      <link>https://arxiv.org/abs/2608.13049</link>
      <guid>https://arxiv.org/abs/2608.13049#2026-08-14</guid>
      <pubDate>Fri, 14 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 OmniScientist: An Omni-Modal Omni-Discipline AI Scientist (2026-08-14)</title>
      <description>Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows, from hypothesis generation and code execution to manuscript preparation. Yet workf...</description>
      <link>https://arxiv.org/abs/2608.13558</link>
      <guid>https://arxiv.org/abs/2608.13558#2026-08-14</guid>
      <pubDate>Fri, 14 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos (2026-08-14)</title>
      <description>Talking-video character replacement requires coordinated transfer of appearance and voice while preserving the source motion, scene, linguistic content, and audio-video timing. Existing methods use se...</description>
      <link>https://arxiv.org/abs/2608.11752</link>
      <guid>https://arxiv.org/abs/2608.11752#2026-08-14</guid>
      <pubDate>Fri, 14 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses (2026-08-13)</title>
      <description>Recent work on distillation transfers the capabilities of large models to smaller ones often by updating the latter's parameters, through teacher forcing, on-policy distillation, and related training-...</description>
      <link>https://arxiv.org/abs/2608.12307</link>
      <guid>https://arxiv.org/abs/2608.12307#2026-08-13</guid>
      <pubDate>Thu, 13 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill (2026-08-13)</title>
      <description>Turning a research idea into a complete paper requires more than text generation: the system must retrieve literature, design and execute experiments, revise claims according to evidence, produce publ...</description>
      <link>https://arxiv.org/abs/2608.11924</link>
      <guid>https://arxiv.org/abs/2608.11924#2026-08-13</guid>
      <pubDate>Thu, 13 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization (2026-08-13)</title>
      <description>Previsualization is an intermediate layer between ideas and production in film, games, architecture, and urban design. It lets creators iteratively refine scenes, actions, cameras, and spatial-tempora...</description>
      <link>https://arxiv.org/abs/2608.12314</link>
      <guid>https://arxiv.org/abs/2608.12314#2026-08-13</guid>
      <pubDate>Thu, 13 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control 📦代码 (2026-08-13)</title>
      <description>LLM-agent services repeatedly execute small deterministic transitions between model and tool calls: route an outcome, update state, and emit the next effect. We ask when this control path exposes enou...</description>
      <link>http://arxiv.org/abs/2608.12123v1</link>
      <guid>http://arxiv.org/abs/2608.12123v1#2026-08-13</guid>
      <pubDate>Thu, 13 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 When the Knowledge Base Becomes the Gold Standard: Measuring Resource-Shared Evaluation Loops in Entity-Level Machine Translation 📦代码 (2026-08-13)</title>
      <description>The Seungjeongwon Ilgi, a UNESCO Memory of the World record, is only 37.4% translated, and the most conspicuous failure mode in automatic translation is the person name -- a misread name corrupts the ...</description>
      <link>http://arxiv.org/abs/2608.11843v1</link>
      <guid>http://arxiv.org/abs/2608.11843v1#2026-08-13</guid>
      <pubDate>Thu, 13 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection (2026-08-13)</title>
      <description>Videos captured through glass often contain reflections that degrade visual quality and interfere with downstream vision tasks. Although single-image reflection removal has been extensively studied, v...</description>
      <link>https://arxiv.org/abs/2608.11562</link>
      <guid>https://arxiv.org/abs/2608.11562#2026-08-13</guid>
      <pubDate>Thu, 13 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 MBA: Multimodal Benchmark and Agents for Real-World Business Ideation (2026-08-13)</title>
      <description>Agentic systems powered by large language models (LLMs) have opened new opportunities for business ideation. Yet existing approaches remain confined to a text-only paradigm, despite the inherently mul...</description>
      <link>https://arxiv.org/abs/2608.11616</link>
      <guid>https://arxiv.org/abs/2608.11616#2026-08-13</guid>
      <pubDate>Thu, 13 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 Persistent Recursive Worlds Enable Autonomous Software Evolution (2026-08-13)</title>
      <description>Complex software systems develop over timescales that exceed the lifespan of any individual coding agent. Most agentic software systems preserve continuity through persistent sessions, memories, manag...</description>
      <link>https://arxiv.org/abs/2608.10450</link>
      <guid>https://arxiv.org/abs/2608.10450#2026-08-13</guid>
      <pubDate>Thu, 13 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 AVA-Encoder: Towards Agent-Native Video Representation Learning (2026-08-13)</title>
      <description>Creative agents still lack an effective way to learn from high-quality human films, limiting their ability to produce cinematic-grade videos. A key challenge is the absence of a structured video repre...</description>
      <link>http://arxiv.org/abs/2608.12313v1</link>
      <guid>http://arxiv.org/abs/2608.12313v1#2026-08-13</guid>
      <pubDate>Thu, 13 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 Constructing Dynamic Master Logic Models as Knowledge Graphs for Complex System Diagnostics Using Retrieval-Augmented Large Language Models (2026-08-13)</title>
      <description>Dynamic Master Logic (DML) provides a hierarchical framework for representing system behavior by linking functional objectives to underlying structural elements. However, DML construction typically re...</description>
      <link>http://arxiv.org/abs/2608.12304v1</link>
      <guid>http://arxiv.org/abs/2608.12304v1#2026-08-13</guid>
      <pubDate>Thu, 13 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 Not a Monolith: Lab-Level Divergence in the Cooperative Equilibria of Chinese Frontier LLM Agents 📦代码 (2026-08-12)</title>
      <description>Does the cooperative bias documented for Western frontier LLM agents extend to a different alignment lineage, and should the Chinese models that embody it be treated as a single bloc or as distinct la...</description>
      <link>http://arxiv.org/abs/2608.10262v1</link>
      <guid>http://arxiv.org/abs/2608.10262v1#2026-08-12</guid>
      <pubDate>Wed, 12 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 The CASE Framework: A Multi-Disciplinary Control Architecture for Governing Enterprise Agentic AI 📦代码 (2026-08-12)</title>
      <description>Enterprises are deploying autonomous AI agents faster than they can govern them, and prevailing approaches stretch a single discipline, typically DevSecOps built for deterministic automation, across e...</description>
      <link>http://arxiv.org/abs/2608.10153v1</link>
      <guid>http://arxiv.org/abs/2608.10153v1#2026-08-12</guid>
      <pubDate>Wed, 12 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence (2026-08-12)</title>
      <description>Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied. We introduce Ex-Omni-2D, an omni-modal dialogue framework th...</description>
      <link>https://arxiv.org/abs/2608.10720</link>
      <guid>https://arxiv.org/abs/2608.10720#2026-08-12</guid>
      <pubDate>Wed, 12 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss (2026-08-12)</title>
      <description>Fréchet distance has recently emerged as an effective distribution-level objective for generator post-training, complementing the conventional sample-level diffusion and flow-matching losses. However,...</description>
      <link>https://arxiv.org/abs/2608.11205</link>
      <guid>https://arxiv.org/abs/2608.11205#2026-08-12</guid>
      <pubDate>Wed, 12 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure (2026-08-12)</title>
      <description>Self-evolving agents accumulate reusable skills by appending successful procedures and failure fixes. Over time, the same requirement is often restated in several branches, examples, and warnings, whi...</description>
      <link>https://arxiv.org/abs/2608.11079</link>
      <guid>https://arxiv.org/abs/2608.11079#2026-08-12</guid>
      <pubDate>Wed, 12 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information (2026-08-12)</title>
      <description>Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scattered across multiple applications (apps) to complete user ins...</description>
      <link>https://arxiv.org/abs/2608.10692</link>
      <guid>https://arxiv.org/abs/2608.10692#2026-08-12</guid>
      <pubDate>Wed, 12 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation (2026-08-12)</title>
      <description>We study reference-free post-training for multilingual machine translation with open large language models. Starting from the supervised-finetuned MiLMMT-46-v0.1 models, we apply Group Relative Policy...</description>
      <link>https://arxiv.org/abs/2608.10812</link>
      <guid>https://arxiv.org/abs/2608.10812#2026-08-12</guid>
      <pubDate>Wed, 12 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation (2026-08-12)</title>
      <description>Visual document retrieval (VDR) is dominated by multi-billion-parameter models that are slow to index at full corpus scale and expensive to serve. Prior compression routes either train a smaller multi...</description>
      <link>https://arxiv.org/abs/2608.10636</link>
      <guid>https://arxiv.org/abs/2608.10636#2026-08-12</guid>
      <pubDate>Wed, 12 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 Beyond Pixels: From Video Priors to 4D Worlds (2026-08-12)</title>
      <description>4D generation synthesizes dynamic 3D scenes from conditions such as text or images. Existing methods either reconstruct generated RGB videos with a separate 4D model or adapt a particular video genera...</description>
      <link>https://arxiv.org/abs/2608.10744</link>
      <guid>https://arxiv.org/abs/2608.10744#2026-08-12</guid>
      <pubDate>Wed, 12 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments? (2026-08-12)</title>
      <description>Real-world data science involves long-horizon workflows that span data wrangling, exploration, modeling, visualization, and validation, and require coordinated use of tools such as notebooks, IDEs, te...</description>
      <link>https://arxiv.org/abs/2608.10366</link>
      <guid>https://arxiv.org/abs/2608.10366#2026-08-12</guid>
      <pubDate>Wed, 12 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring (2026-08-11)</title>
      <description>As AI coding agents take on increasingly complex, long-horizon software engineering tasks, existing benchmarks are rapidly saturating and their evaluation quality has come under serious scrutiny: a re...</description>
      <link>https://arxiv.org/abs/2608.09802</link>
      <guid>https://arxiv.org/abs/2608.09802#2026-08-11</guid>
      <pubDate>Tue, 11 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA (2026-08-11)</title>
      <description>Macaron-V1 is an open agent-model family for experiential intelligence: learning from experience in real environments and continuing to learn after deployment. It is organized around two system goals....</description>
      <link>https://arxiv.org/abs/2608.09819</link>
      <guid>https://arxiv.org/abs/2608.09819#2026-08-11</guid>
      <pubDate>Tue, 11 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains (2026-08-11)</title>
      <description>We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains. It contains 1,253 expert-annotated examples spanning 60...</description>
      <link>https://arxiv.org/abs/2608.09873</link>
      <guid>https://arxiv.org/abs/2608.09873#2026-08-11</guid>
      <pubDate>Tue, 11 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 Motif 3: Technical Report (2026-08-11)</title>
      <description>We introduce Motif 3, a decoder-only Mixture-of-Experts language model with 314 billion total parameters and 13.2 billion activated per token. Each sparse MoE layer contains 384 routed experts, with e...</description>
      <link>https://arxiv.org/abs/2608.09119</link>
      <guid>https://arxiv.org/abs/2608.09119#2026-08-11</guid>
      <pubDate>Tue, 11 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 SHE: Trajectory-driven Safety Harness Evolution for LLM Agents 📦代码 (2026-08-11)</title>
      <description>The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, permissions, and runtime control. Existing safety m...</description>
      <link>http://arxiv.org/abs/2608.09885v1</link>
      <guid>http://arxiv.org/abs/2608.09885v1#2026-08-11</guid>
      <pubDate>Tue, 11 Aug 2026 18:00:00 +0000</pubDate>
    </item>
    <item>
      <title>📄 Multi-Agent AI Safety as an Institutional Design Problem 📦代码 (2026-08-11)</title>
      <description>AI agents increasingly work inside systems that govern how they delegate tasks, move information, execute actions, and use shared resources. Recent work already shows that deployment rules can change ...</description>
      <link>http://arxiv.org/abs/2608.09828v1</link>
      <guid>http://arxiv.org/abs/2608.09828v1#2026-08-11</guid>
      <pubDate>Tue, 11 Aug 2026 18:00:00 +0000</pubDate>
    </item>
  </channel>
</rss>
