Survey · 2026 · preprint coming soon

From Trajectories to Experience:

A Survey of Experience-Driven LLM Agents

Xiaohui Yan1, Yingchen Zhang2, Shiyao Liu2, Xiaofei Huang2, Yuxin Yue2, Zaiyu Xia2, Jiayu Yao2, Xinjie Chen2, Yun Hu1, FuRong Li1, Shenghua Liu2, Yixing Fan2, Zixuan Li2, Ruqing Zhang2, Jiafeng Guo2

1Huawei Technologies Co., Ltd  ·  2Institute of Computing Technology, Chinese Academy of Sciences

How can an LLM agent turn its trajectories into guidance that improves later behavior, and how can we tell when it does?

At a Glance

LLM agents record in their trajectories which strategies failed and how they recovered, yet an agent that resolves a problem in one episode may repeat the same diagnostic detour in the next. Work on reflection, agent memory, workflow learning, and skill acquisition addresses this gap under different names: a distilled lesson is a self-reflection in Reflexion, an insight in ExpeL, a causal abstraction in CLIN, and a memory item in ReasoningBank. Because names and content do not line up, it is hard to see which guidance a method retains, when that guidance applies, and how to compare reported gains.

The survey makes four contributions.

1

A unified framework

Every experience unit, whether a precedent, a lesson, or a procedure, is described by its guidance, applicability conditions, execution evidence, and reliability evaluation. Mechanisms are organized by four lifecycle functions: formation, organization, utilization, and maintenance.

2

Evidence standards

Evaluation of intrinsic experience quality, the experience utilization process, and downstream effects is connected to three improvement claims: fixed experience reuse, feedback-driven sustained improvement, and improvement of the learning mechanism.

3

Application analysis

Seven application settings are related to four conditions: how often tasks recur, how quickly feedback arrives, whether outcomes can be checked, and how readily environmental changes invalidate earlier guidance.

4

Research agenda

Eight open problems, from learning with human expert records to correcting propagated errors, aimed at improvement that is cumulative, transferable, and correctable.

From LLMs to Experience-Driven Agents

Three capability emphases: Era of LLM (Know), Era of Agent (Know + Act), Era of Experience (Know + Act + Evolve).
Three capability emphases. Each keeps the capabilities of the previous one and adds to them; the experience loop is the focus of the survey.

Era of LLM · Know

A knowledge-centric paradigm: pretrained models encode knowledge, retrieval-augmented generation supplies more, and chain-of-thought helps models reason with it.

Era of Agent · Know + Act

An execution-centric paradigm built on planning, tool use, and runtime feedback. Correcting an error within one attempt, however, does not automatically improve the next.

Era of Experience · Know + Act + Evolve

An experience-centric paradigm: retain what trajectories reveal in an explicitly accessible form, use it to guide later episodes, and revise it as feedback accumulates.

An experienceless agent repeats an import error; an experience-driven agent verifies its fix, stores the lesson, and reuses it in the next episode.
An experienceless agent versus an experience-driven agent. After verifying a fix, the experience-driven agent stores “on import errors, check dependencies” in its experience repository and reuses it in episode k+1.

Agents differ in whether their retained experience changes during later work. In fixed-repository reuse the repository stays fixed; in closed-loop learning later trajectories inform the formation and maintenance of experience, and the updated repository guides subsequent episodes. This cycle is the experience loop. One success supports a lesson only in comparable situations: another import failure may have a different cause.

Foundations

What counts as agent experience

Agent experience is action-guiding information that is grounded in one or more task-execution trajectories and retained to guide later episodes where it applies.

Execution grounding

The information derives from task-execution trajectories: the agent's own, other agents', or human practice, including skills people author from their own work.

Cross-episode retention

It is kept beyond the episode that produced it. A correction discarded together with the attempt is not experience.

Behavioral relevance

Its role is to shape later behavior: interpreting a question, producing an answer, evaluating a solution, or acting in an environment.

Correctness and benefit are not required: later checks test them. Outside the core scope are audit-only logs, factual or preference memory without a guidance role, instructions with no task practice behind them (generic instructions, tool documentation), and learning solely through parameter updates.

Three representational forms

Consider one code-repair episode: tests fail at the import stage, the agent traces the failure to a declared dependency missing from the active environment, restores it, and confirms that the import succeeds. The agent can retain three forms of experience from it.

Trajectory-level · precedent

Retained: a diagnostic example — the import error, relevant environment details, dependency checks, repair action, and successful import check.

Guides by: showing how a comparable failure was diagnosed and fixed.

e.g. RAP, Synapse, JARVIS-1

Semantic-level · lesson

Retained: “When tests fail during import in a comparable setup, check declared dependencies and the active environment before changing application logic.”

Guides by: prioritizing an environment check without assuming the new failure has the same cause.

e.g. Reflexion, ExpeL, AutoGuide

Procedural-level · procedure

Retained: inspect the error; compare declared dependencies with the active environment; resolve a confirmed mismatch within task permissions; rerun the import check.

Guides by: ordered steps, with current observations deciding whether a repair is warranted.

e.g. Voyager, AWM, SkillWeaver

Form describes only how guidance is expressed. The levels are not stages of maturity and do not rank generality: a scoped lesson can apply as narrowly as the precedent it came from, and several forms can coexist. All three have developed side by side since 2023.

Representative experience-driven agents grouped into trajectory-level, semantic-level, and procedural-level branches by year from 2023 to 2026.
Representative experience-driven agents by principal form and year of first release. Branches group methods by form and do not indicate lineage.

The experience unit ⟨g, c, e, r⟩

Whatever its form, every piece of experience — an experience unit — can be described along four analytical dimensions. They are questions to ask of any unit, not fields a system must store.

Ei = ⟨ gi, ci, ei, ri ⟩

g

Guidance

What the unit tells the agent to do, avoid, check, or consult.

Check declared dependencies and the active environment first.

c

Applicability conditions

Task, environment, tool, and executor assumptions under which it applies.

An import-stage failure, a comparable setup, permission to inspect the environment.

e

Execution evidence

The trajectories, or parts of them, that support it, and any counterexamples.

The error, the inspection, the repair, the passing import check.

r

Reliability evaluation

How strong that support is and how uncertain it remains.

Provisional after a single episode.

Reliability asks whether guidance is well supported under its stated conditions. Utility asks whether using it pays off for a particular executor and task; it depends on context and is not a fixed attribute of the unit. If the executor already checks dependencies unprompted, the lesson may be reliable yet add little utility.

Four experience types

Independently of form, experience can be classified by the execution outcome its guidance draws on. Each type needs a different check before its guidance is relied on.

TypeEvidence drawn onMain cautionExamples
SuccessA successful trajectorySuccess does not show which steps were necessary, and its label may be wrong.AWM, Voyager
FailureA failed trajectory, with its feedbackA diagnosis that blames the wrong step yields an overly broad prohibition.Reflexion, ReasoningBank
ComparativeSuccessful and failed attempts at comparable tasksUnless attempts are aligned, a difference may reflect the environment or task difficulty.ExpeL, AutoGuide
CorrectiveAn error linked to the change that resolved itThe same error may have a different cause in a new setting.EXG

The experience loop

During episode k, retained experience is turned into situated guidance for the current decision:

Ak,t = Utilize(Rk, qk, hk,≤t)

Closed-loop learning adds a path from execution back to the experience repository R:

(yk, Tk) = Execute(qk, Rk)
𝓔k = Form(Tk, Rk, 𝓗<k)
Rk+1 = Update(Rk, 𝓔k, Tk)

Formation proposes a possibly empty set of candidate units 𝓔k; maintenance decides which enter Rk+1 and how retained units change. An experienceless agent never receives situated guidance, fixed-repository reuse keeps Rk unchanged, and producing Ak,t does not by itself retain a new unit.

The Experience Lifecycle Framework

Making the experience loop work requires four kinds of decisions: what to learn from a trajectory, how to store it so that it can be found, how to use it to guide behavior or train the executor, and what to keep or change as evidence accumulates.

The experience lifecycle: formation derives candidates, maintenance and organization manage the experience repository, and utilization supplies guidance for subsequent execution.
The agent experience lifecycle, illustrated through runtime reuse. Training-time internalization provides another utilization route.
FunctionCore questionMain responsibility
Experience formationWhat can be learned from past trajectories?Select or derive candidate precedents, lessons, or procedures, and check that the trajectories support them.
Experience organizationHow should experience be stored so that it can be found and interpreted?Index experience and link it to its use contexts, applicability conditions, and related units.
Experience utilizationHow should experience guide behavior or train the executor?Decide when to seek experience, which units to use, and how to adapt them; in training, use experience to shape trajectories and supervision.
Experience maintenanceWhat should be admitted, kept, revised, or retired?Admit candidates, and revise, consolidate, or retire retained experience as evidence and conditions change.

The division also locates failures. For the dependency lesson, a version claiming that every import failure requires installing a dependency points to unsupported generalization during formation; a retained lesson missing from the lookup index points to organization; applying it despite incompatible environment assumptions points to utilization; and leaving it unchanged after new evidence contradicts it points to maintenance.

Experience Formation

Browse 43 papers →

Formation turns trajectories into candidate experience: a precedent to consult, a lesson to apply, or a procedure to follow. The survey's central argument here: the more a mechanism adds beyond what the trajectories recorded, the more its support must come from outside them — through counterexamples, tests beyond the source tasks, or new execution.

Formation mechanisms for trajectory-level, semantic-level, and procedural-level experience.

Trajectory-level

  • Select and annotate RAP, SEER, TRAD
  • Segment and compress Synapse
  • Relabel and reconstruct ECHO, BAGEL, BREW

Semantic-level

  • Reflective abstraction Reflexion, CLIN, ReasoningBank
  • Comparison and diagnosis AutoGuide, ExpeL, TF-GRPO
  • Cross-trajectory generalization AutoManual, EMG, EDV

Procedural-level

  • Extraction and parameterization AWM, SSO, TraceCompiler
  • Synthesis and composition Voyager, SkillWeaver, Metis
  • Execution-guided verification ASI, SkillCAT, SkillOpt

Key findings

  • Deciding more in advance pays off only while conditions stay stable. A precedent leaves interpretation to the receiving agent, a lesson states a conclusion, and a procedure fixes the steps. In Metis (AppWorld, GPT-4o executor), when each task could use only memory from earlier tasks, text memory reached 73.3% task success against 53.3% for code.
  • No form wins on every benchmark. In ExpeL, insights beat retrieved trajectories on HotpotQA (36% vs 31%) while trajectories win on ALFWorld (55% vs 50%); in Memp on ALFWorld with GPT-4o, scripts alone reach 56.4%, trajectories 74.3%, and both 77.9%. Complementary forms help.
  • Each formation check supports only what it tests. Source success concerns the recorded behavior, a model's judgment remains an interpretation, a successful rerun shows the candidate works under the tested conditions, and only held-out tasks test transfer.

Experience Organization

Browse 47 papers →

A stored unit helps only if the agent can find it and reach the context needed to read it. Even a correct match leaves four questions: Q1 was the fix recorded with the failure? Q2 is the condition under which it worked kept with it? Q3 is the match one step of a longer procedure, or related to other precedents? Q4 can the agent open what lies behind or inside the match?

Flat, tree-structured, graph-structured, and hierarchical-graph experience organization.
PatternUseful whenUpkeep and failure modes
FlatA symptom, task, or usage description matches the request, and the record itself holds the repair and its conditions (Q1–Q2).Keys must anticipate later needs; a missing key fails silently and relations stay implicit.
TreeMany similar records call for a narrower place to search, or the missing context is a broader group, containing routine, or base procedure (Q3).Branch choices can omit records; revising a shared base affects every variant.
GraphThe repair, a continuation, or a supporting skill sits in another record that would not match on its own (Q1, Q3).Relations must be built, checked, and updated; unreliable links and wide expansion mislead or overload access.
Hierarchical graphThe match needs its source trajectory or implementation, or must be checked or trimmed before use (Q4).Cross-layer correspondences must stay consistent; expansion adds inspection cost.

Key findings

  • Better keys before added structure. The right record is often missed only because of how it is keyed; keying it by the observable symptom, as MemGovern's repair cards do with error signatures, is the cheaper first remedy. Added structure should beat that improved-key baseline.
  • Structure keeps conditions, evidence, and reliability reachable — condition keys make applicability searchable and source links keep supporting trajectories reachable — but the agent must still judge whether the guidance fits.
  • Reported ablations rarely isolate structure from content. Removing a level or an edge usually removes what it holds, so whether structure itself explains reported gains remains open.

Experience Utilization

Browse 60 papers →

An agent holding useful experience may still overlook it, follow a familiar procedure under the wrong conditions, or spend more effort consulting experience than the task requires. At runtime, utilization answers four connected questions; during training, experience can also be internalized into model parameters.

Experience utilization: triggering, selection, adaptation, and application, with paths for direct reuse and continuing without experience.

Experience triggering

When should the agent consult experience?

Fixed schedules, observed failures, internal-state signals, coach gates, learned policies.

Experience selection

Which experience, if any, is worth using?

Retrieval, reranking and filtering, and allocation across multiple agents.

Experience adaptation

What should change before reuse?

Contextual grounding, guidance specialization, procedural reconfiguration, feedback-guided revision.

Experience application

How should guidance shape behavior?

Five functions: interpretation, generation, decision-making, procedure reuse, and evaluation.

Experience internalization transfers experience-supported behavior into model or adapter parameters: training on experience-shaped trajectories, distilling experience-conditioned policies, or withdrawing guidance during training.

Key findings

  • The four decisions share one budget. Seeking, selecting, and adapting experience all cost resources before the guidance helps, so inexpensive checks that could change the reuse decision belong before costly actions.
  • Recurrence decides whether experience is worth it. On largely self-contained web tasks, giving an executor without experience more interaction steps can match or exceed online skill and memory augmentation at approximately matched token cost.
  • Adapting before reuse pays. With DeepSeek-V4-Pro, QCR's query-conditioned notes raise average success by 10.7 points over the full trajectory while using 48.9% fewer online tokens.
  • Check outcomes independently of the explanation experience proposed; when the same experience shapes both the action and its check, the check can inherit the action's mistaken premise.

Experience Maintenance

Browse 68 papers →

Each episode leaves new candidates and evidence about experience already retained. Simply appending candidates accumulates noise alongside useful guidance; maintenance decides what to admit, combine, correct, or retire so that later reuse stays reliable and efficient.

Experience maintenance: admission, consolidation, refinement, and retirement acting on a shared experience repository.
OperationProblem signalMaintenance choiceMain risk
AdmissionA new candidate is proposedAccept supported value; defer uncertainty; reject unsupported contentExcluding rare useful guidance or accepting a false explanation
ConsolidationEquivalent or fragmented guidance accumulatesMerge equivalents; abstract or share structure while keeping distinctionsCompression removes an action-changing condition
RefinementContradictory outcomes or changed prerequisitesDiagnose; adjust reliability, scope, or actions; validate before replacementRepairing the wrong cause or breaking valid behavior
RetirementSuspected loss of added benefitDownweight, archive, or delete according to evidence and recoverabilityConfusing poor access with low value; losing a rare recovery
CoordinationRepeated formation defects or mismatched accessChange operation selection, formation, or retrievalDelayed rewards obscure which change helped
SafeguardsA premise loses support or contamination is detectedRestrict affected guidance; trace, repair, and validate successorsMissing descendants or removing independently valid guidance

Evidence for these decisions comes at three check levels: a support check examines content against recorded evidence; a paired comparison measures the local gain from using it; a collection-level test evaluates it together with the experience it would join.

Key findings

  • Admission quality can decide whether updating helps at all. In the Experience-Following study, adding every trajectory ends below a fixed copy of the starting memory, while selective addition with a strict ground-truth evaluator ends above it — but deployed agents rarely have such labels.
  • Updates can reverse earlier gains and spread. Consolidation can drop an action-changing detail, and a mistaken update can propagate into experience derived from it, so maintenance must also test unaffected behavior and keep what is needed to repair descendants.
  • Open questions: how broadly to update, which interactions a collection-level test must cover, how to value units whose usage data earlier decisions have filtered, and where to spend verification effort.

Evaluation: Three Targets, Three Claims

Task success alone does not show whether stored experience is correct or whether the agent benefited from using it: an agent may ignore a useful lesson, complete a task despite errors in a stored procedure, or score higher simply because of additional model calls.

Three evaluation targets: intrinsic experience quality, the experience utilization process, and downstream effects with experience frozen or updates enabled.

Intrinsic experience quality

Is retained guidance reliable?

Audits units and the repository against references fixed in advance.

Experience utilization process

Is it used appropriately?

Scores triggering, selection, adaptation, and experience application, and attributes behavior to supplied experience.

Downstream effects

Does it improve later tasks?

Measures effectiveness, cost, transfer, and harm with experience frozen or with updates enabled.

Higher performance with experience does not by itself show that an agent keeps improving. Each improvement claim needs more demanding evidence than the one before:

Improvement claimQuestionNecessary evidence
Fixed experience reuseDoes a fixed body of experience help on new tasks?A matched access/no-access comparison with a frozen repository and a fixed executor, reporting gains and harms under stated resources. For internalization: matched training with and without the frozen experience from the same initial policy, tested without the guidance.
Feedback-driven sustained improvementDo updates from new trajectories keep adding capability?Evidence that multiple updates add capability over a predeclared interval of comparable tasks, against a copy that keeps its starting repository fixed; acquisition, retention, regressions, costs, and uncertainty reported together.
Improvement of the learning mechanismDoes revising how experience is formed or updated improve later learning?Identified, persistent revisions compared with the original mechanism on subsequent learning under comparable tasks and budgets.

What the evidence supports today. Fixed experience reuse has the clearest support, chiefly where tests or reference answers check outcomes; budget-matched controls are uncommon. Evidence for sustained improvement is partial: some studies compare updating streams with a fixed starting memory, but acquisition, retention, regressions, costs, and uncertainty are rarely examined together. No reviewed protocol yet links audited unit quality to downstream gains and harms on the same cases, or checks whether a unit's reliability evaluation matches its evidence.

Browse 43 evaluation papers and benchmarks →

Applications

Seven application settings, compared on the four reuse conditions that determine what experience is worth retaining and how it can be checked.

ApplicationRetained experienceRecurrenceEnvironmental changeOutcome checkingFeedback timing
Question answeringExpeL, Memento, Dynamic Cheatsheet, MERITSolution templates, search lessons, tool-use proceduresReasoning and search difficulties recur across questionsFacts, sources, and databases differ across questionsReference answers; report quality is harder to checkPer question, from answers or query results
Software engineeringSWE-Exp, SWE-ContextBench, EnvPilotDiagnostic precedents, modifications, setup strategiesSimilar symptoms recur, often with different causesVersions, dependencies, and languagesExecutable tests; they do not validate transfer of a diagnosisPer attempt, from tests or setup runs
Computer useAWM, SkillWeaver, MAGNET, CUA-SkillWorkflows, callable and authored skills, interaction correctionsRepeated operations on the same sites and applicationsInterface updates and professional rolesState and artifact checks; clicks alone are insufficientPer step, from interface state
Recommender systemsCRAVE, MemoCRS, Re2LLM, SAGERShared selection lessons, error-correcting hints, per-user decision principlesRecurring requests with varying preferencesPreferences drift and differ across usersOffline, against recorded user choicesPer recommendation, from the user's response
Embodied controlVoyager, CLIN, ViReSkill, Pragmatist RobotEnvironmental lessons, plans, executable skillsFamiliar instructions under different statesObjects, resources, and initial statesSimulator success; physical evidence is limitedAfter each attempt, from task outcome
Scientific researchDS-Agent, EvoScientistSuccessful solutions, direction experience, experimental strategiesModeling tasks recur; research directions varyTask distributions and executor search proceduresRun status is checkable but budget-dependent; research value is notDelayed and graded: run, baseline, evaluation
Trading and forecastingTradingGroup, FinCon, Live-EvoPrior decisions, outcomes, strategy lessonsSimilar market or question conditionsMarket conditions shiftBacktests or resolved forecasts, not prospective returnsFinal outcomes delayed; applicability changes can be checked earlier

Four design decisions

Whether to invest

Build experience only when expected reuse repays the cost of acquiring, checking, and maintaining it. Repeated operations on a stable website can justify tested functions; a rarely recurring question type may justify only a lightweight precedent.

What to retain

Retain the conditions that can be recorded and checked with the guidance: repair stage and dependency context in software, the intended artifact and role in workplace procedures, the failing stage, resources, and state for negative scientific or robot experience.

How strongly to use it

Uncertain experience enters as a candidate explanation, planning option, or checklist. A tested procedure can execute when its inputs, starting state, and required output can be checked; current observations still settle the present request.

When to revise

Trigger review by observed changes, not age: an interface update, new output requirements, a role switch, or worsening outcomes on comparable tasks. With delayed outcomes, revisions justified by the final result wait until it resolves.

Browse 31 application papers →

Main Conclusions

  1. Keep guidance connected to its conditions and evidence.

    The central design principle: throughout formation, reuse, and revision, keep guidance tied to its applicability conditions and to evidence that can support or challenge it. A more compact or executable representation need not preserve clearer conditions or stronger support.

  2. Condensed guidance can lose the conditions that made it work.

    “On import errors, check dependencies,” condensed further into “reinstall dependencies,” skips the step that confirms a dependency is missing. Keeping the source trajectory with the lesson makes such omissions recoverable.

  3. Shared experience must fit the agent that receives it.

    LEGOMem's ablations show that where task and subtask memories are placed affects success, and a more capable teacher does not always produce the most helpful memory for a given student.

  4. Entries harmless individually can be harmful together.

    Entries that a model-based auditor mostly judges benign separately can elicit harmful behavior when a query activates them together, so units must be tested in the combinations later queries activate.

  5. Rising task-stream scores mix learning with changing task demands.

    Separating the two requires comparing the updating agent with a copy whose starting repository stays frozen, on comparable tasks.

  6. Fixed experience reuse is the best-supported claim.

    Evidence is clearest where tests or reference answers check outcomes. Sustained improvement is supported only partially, and improvement of the learning mechanism further requires testing the revised mechanism's contribution to subsequent learning.

  7. Reuse shapes what the agent learns next.

    A familiar strategy may suppress exploration, and a mistaken premise may propagate into later experience. Preserving sources and conditions lets later evidence challenge earlier guidance.

Limits: the survey reviews a selective, mechanism-oriented corpus rather than a systematic screen; much evidence comes from recent preprints; assigning evidence roles and improvement claims is the authors' reading of each experimental design; and the proposed evaluation protocols have not yet been implemented.

Eight Open Problems

Toward improvement that is cumulative, transferable, and correctable. Each problem pairs why it remains open with what would count as progress.

Acquiring and extending experience

1

Forming experience from human expert records

Open because records omit rejected alternatives, reasons, and environment assumptions, and reflect tools and permissions the agent may lack.

Progress: human records reduce the interactions needed for a given success rate, counting human effort, without importing assumptions the agent cannot meet.

2

Preserving informative exploration

Open because successes under a dominant strategy cannot show whether an alternative would also have worked, and an abandoned strategy may never be retried.

Progress: the agent notices when a once-effective lesson has become too restrictive, at lower interaction cost than exploring only after failures.

3

Transferring principles across domains

Open because the base model may already state the general principle; knowing when it applies and when it does not is the hard part.

Progress: on held-out domains, the agent decides when to apply or withhold a principle more accurately than with domain-specific lessons or the base model alone.

4

Composing experience for unseen tasks

Open because two procedures can each succeed alone and fail together; a stored link cannot show that one produces what the next requires.

Progress: the agent solves unseen combinations of known procedures and detects incompatible ones with limited checking.

Keeping experience useful, revisable, and safe

5

Co-evolving experience representations and executors

Open because as the executor learns, or across executors sharing a repository, some guidance becomes redundant and some must change form.

Progress: the executor drops support it no longer needs while still handling exceptions and learning from new experience.

6

Tracing and correcting propagated errors

Open because a flawed premise can pass into later experience without its wording or a recorded link, and privacy limits how much history can be kept.

Progress: the agent removes an inherited flawed premise despite incomplete records, without discarding behavior that has independent support.

7

Learning under adversarial interaction

Open because an adversary can craft experience that passes checks on each unit, spread harm across units, or delay its effect.

Progress: defenses limit adaptive attacks under stated access assumptions while losing little learning on benign tasks.

Evaluating cumulative progress

8

Establishing cumulative capability gains

Open because rankings of learning strategies can reverse when later tasks reward different capabilities, so one task stream cannot settle them.

Progress: gains hold across successive snapshots of the agent, reported with regressions and total cost, with final tests kept out of update decisions.

Cite

If the survey or this list helps your research, please cite:

@misc{yan2026trajectories, …}

Cite this survey

If this list helps your research, please cite From Trajectories to Experience: A Survey of Experience-Driven LLM Agents.