Code-based skills for long-horizon language agents

Up and Down the Abstraction Ladder

We study the benefits and tradeoffs of using low-level primitives, higher-level code skills, or both for long-horizon language agents.

Bartłomiej Cupiał1,61 University of Warsaw
6 AKCES NCBR
, Jens Tuyls22 Princeton University, Maciej Wolczyk33 IDEAS NCBR, Davide Paglieri44 University College London, Martin Klissarov5,75 McGill University
7 Mila
, Benjamin Eysenbach22 Princeton University, Piotr Miłoś1,8,91 University of Warsaw
8 Mistral AI
9 Institute of Mathematics, Polish Academy of Sciences
, Karthik R. Narasimhan22 Princeton University

Rollouts

Watch what the agents play.

01

GPT-5 mixed-control trajectory in NetHack

8 decisions / sec
03

When skills are insufficient, the agent moves down the abstraction ladder.

Two examples show primitive actions acting as an escape hatch: first after a skill fails, and then when the required behavior is missing from the skill library.

Skill Primitives Skill
Failed skill Primitive recovery 9 sec

Recovering after ranged combat fails

Ranged combat runs out of ammunition. GPT-5 switches to primitive inputs, kills the gas spore, then returns to a high-level pickup skill.

Missing skill Primitive fallback 13 sec

Filling the drop-gold skill gap

The vault guard demands the gold, but no drop-gold skill exists. GPT-5 uses primitive inputs to drop all gold, then returns to exploration.

Abstract

In NetHack and MiniHack, code skills improve progression, reduce inference cost, and accelerate learning.

Acting and learning in long-horizon environments remains challenging for language agents, especially on tasks requiring low-level action control such as games. Primitive action control often provides poor grounding for language models and, when combined with long horizons, can be prohibitively expensive and hinder learning, exploration, and planning.

Inspired by work on agent interfaces and horizon reduction, we study the value of using higher-level, semantically meaningful code skills in addition to or instead of low-level primitives. We develop CodeHack, a rich library of code-based skills with natural-language descriptions, and compare agents that can only use primitive actions, only use semantic skills, or use both.

We evaluate these agents in three settings: zero-shot prompting, supervised fine-tuning, and reinforcement learning. Averaged across a broad zero-shot NetHack evaluation, skills provide over 3x improvement in game progression while cutting inference costs per episode by 86%. Mixed control retains much of this benefit while preserving a path back down to low-level actions, and in RL, skill-based agents learn significantly faster than agents acting on primitives.

Skills improve NetHack performance in zero-shot use and during reinforcement learning

Skills vs. primitives in NetHack

Higher-level, semantically meaningful skills provided by CodeHack improve performance both in zero-shot use and during RL. Learning curves are for Qwen-3.5-4B.

3.4x
Higher zero-shot NetHack progression
Skill-only vs. primitive-only, averaged across the 14-model NetHack sweep.
86%
Lower estimated inference cost
Per-episode cost reduction for skill-only control in the same zero-shot sweep.
5.1x
Larger RL dungeon-level gain
Skill-only vs. primitive-only across the two open controllers at the fixed PPO budget.

Motivation

Why climb the abstraction ladder?

As language models transition into autonomous agents, they are increasingly tasked with complex, long-horizon problems, ranging from software engineering and computer use to embodied control. To succeed in these environments, agents must manage extended interactions, adapt to shifting contexts, and recover from inevitable errors, all while carefully managing their inference budgets.

Despite this, most agents are still forced to operate through low-level action interfaces, expanding simple behaviors into long sequences of brittle, error-prone decisions. Agentic coding systems make this tension concrete: if a coding agent had to generate one keystroke at a time, it would risk wasting its reasoning budget on mechanics rather than design. In practice, programmers and coding agents rely on functions, libraries, tools, and editor commands that package many low-level operations into semantically meaningful units.

This suggests a central question for long-horizon language agents: what are the benefits and tradeoffs of using primitive actions, higher-level skills, or both? We study this question through temporally extended code-based actions, or skills. Code-based skills are natural for language agents: they can be described in language, inspected as source code, executed cheaply on the CPU, and combined with primitive actions when finer control is needed.

We use NetHack as our main domain because it captures much of the structure of real long-horizon agentic tasks: agents must act under partial observability, adapt across changing contexts, recover from mistakes, and repeatedly switch among qualitatively different modes of behavior such as exploration, navigation, combat, resource management, and tool use. This complexity also makes pure skill abstraction incomplete in practice, motivating agents that can move flexibly up and down the abstraction ladder.

Method

CodeHack turns one skill call into many game actions.

CodeHack is a library of Python skills that sits between a controller and the NetHack or MiniHack environment. Each skill maps the current symbolic game state to a sequence of primitive commands, such as movement, inventory interactions, prompt responses, or combat actions. The controller still makes the high-level decision, but repeated low-level mechanics can run through cheap, inspectable code.

The same runtime exposes three action interfaces: primitive-only, skill-only, and mixed. This lets the paper compare abstraction levels while keeping the environment and evaluation protocol fixed. Skills return control when they finish, fail, trigger a panic handler such as damage or a newly reachable hostile monster, or make no progress, giving the outer controller a chance to choose another action.

CodeHack as an intermediate control layer between controllers and NetHack or MiniHack environments
CodeHack is an intermediate control layer between a controller and NetHack or MiniHack. The controller selects either an abstract code skill, such as explore or fight_melee, or a low-level primitive action. The runtime maintains symbolic state, inventory tracking, map memory, pathfinding, panic handlers, and no-progress feedback while expanding skills into primitive environment steps.

Semantic actions

Skills correspond to recognizable behaviors such as exploration, combat, equipment management, and terrain interaction.

Reusable code skills

Skills are reusable across many states and episodes, functioning as stable units of abstraction rather than narrow scripts.

Mixed abstraction levels

The interface supports moving up and down the ladder, so skills accelerate long-horizon behavior without removing local intervention.

Experiments

We test how skills affect zero-shot performance, inference cost, mixed control, and RL.

The experiments evaluate the same action-interface choices in zero-shot prompting, supervised fine-tuning, and reinforcement learning. MiniHack provides controlled tasks for navigation, exploration, combat, item use, and subgoal sequencing, while NetHackScore-v0 serves as the main long-horizon testbed. The paper tracks MiniHack success rate and NetHack score, progression, dungeon depth, token use, and estimated inference cost.

Zero-Shot Gains

Skills improve task success in MiniHack and game progress in NetHack.

With GPT-5 on MiniHack, CodeHack skills improve success on every task and produce a 55 percentage-point average absolute gain over primitive control. Quest-Hard is the sharpest example: skill control reaches 25% success, while primitive control solves no episodes. In the broader NetHack sweep across 14 models and five model families, skill-only agents achieve 3.4x higher progression, 4.3x higher score, and about 2.7x deeper dungeon reach than primitive-only agents.

Zero-shot MiniHack and NetHack results comparing skills and primitives

Zero-shot results across MiniHack and NetHack

Left: MiniHack success rates for GPT-5. Right: NetHack milestone reach for GPT-5. Both plots present averages over 32 episodes per task.

Performance-Cost Frontier

Skills reach deeper dungeon levels at lower or comparable cost.

Skills improve the frontier by replacing many expensive language-model decisions with CPU execution inside the CodeHack runtime. Across the zero-shot NetHack sweep, skill-based agents generally shift up and to the left: they reach deeper dungeon levels while using lower or comparable per-episode cost. Skill-only control reduces estimated cost by 86% and token use by 70% compared with primitives, while mixed control remains far stronger and cheaper than primitive-only play.

Cost-performance frontier plotting dungeon level against inference cost

Zero-shot NetHack cost-performance frontier

Each point corresponds to a model-interface pair and plots average dungeon level reached against average inference cost per episode. Skill-based agents define the strongest frontier, while mixed control is typically intermediate.

Faster Learning

RL amplifies the performance gap between skill-based control and primitive-only control.

PPO training on Llama-3.1-8B-Instruct and Qwen-3.5-4B shows that abstraction helps learned controllers too. Under the same training budget, skill-only control produces a 5.1x larger gain in dungeon level than primitive-only control, and mixed control reaches a 6.5x larger gain. Supervised fine-tuning on teacher trajectories followed by RL produces the strongest learned skill controllers reported in the paper, but the central pattern is already visible from RL alone: shorter effective horizons make learning faster.

RL training curves comparing skills, primitives, and mixed control

RL training curves for Qwen-3.5-4B

Training curves compare action interfaces over the same PPO budget. Skill-based and mixed interfaces improve progression and dungeon depth faster than primitive-only control.

Citation

Cite this work.

@misc{cupial2026abstractionladder,
  title = {Up and Down the Abstraction Ladder: Code-Based Skills for Language Agents},
  author = {Cupial, Bartlomiej and Tuyls, Jens and Wolczyk, Maciej and Paglieri, Davide and Klissarov, Martin and Eysenbach, Benjamin and Milos, Piotr and Narasimhan, Karthik R.},
  year = {2026},
  note = {Preprint}
}