When you first dive into reinforcement learning, it is tempting to think that the magic lies almost entirely in the algorithm. We obsess over policy gradients, reward scaling, and model architecture. But the more you work on training agents to operate in complex domains—whether that means navigating a codebase or sculpting a scene in Blender—the more you realize a stark truth: an agent’s capability is fundamentally bounded by the world you build for it.

Recently, as Large Language Models have started pushing past the static limits of supervised pre-training, there’s been a massive revival in RL environments. People want models that can explore outside their training distribution, figure out novel pathways, and use tools iteratively. But building a proper “RL harness” is an art form—one that sits right at the intersection of system design, math, and subtle software engineering.

Here are a few core insights and lessons learned on how to approach environment design, why simple intuition often breaks down, and what it actually takes to train agents that learn.


The Illusions of Perception: Observations Aren’t States

In textbook RL, we often pretend the agent looks out at the world and sees everything exactly as it is. In practice, real environments almost never work this way.

An agent doesn’t receive the environment’s true internal state; it receives an observation—a partial, noisy, and often redundant glimpse of reality. If you give an agent a single snapshot of a scene, it loses all concept of velocity, acceleration, and intent. To make intelligent decisions, the agent needs a sequence of observations over time to infer the underlying trajectory.

       ┌────────────────────────┐
       │   Environment State    │ (Hidden / Complete)
       └───────────┬────────────┘
                   │
           [ Filter / Mask ]
                   │
                   ▼
┌──────────────────────────────────────┐
│  Observation Stream: (o_t-2, o_t-1)  │ ──► Agent Decision
└──────────────────────────────────────┘

This applies to time bounds as well. If an agent has a deadline, it needs to know how much time is left. A common pattern in environment design is encoding cyclical variables—like remaining execution windows or continuous physical loops—into smooth sine and cosine embeddings.

A Quick Tangent: Why Transformers Love Sines and Cosines

This idea of passing cyclical time to an agent mirrors how Transformers handle sequence order. When you add fixed positional encodings to token embeddings:

\[PE_{(pos, 2i)} = \sin\left(\frac{pos}{10000^{2i/d_{\text{model}}}}\right), \quad PE_{(pos, 2i+1)} = \cos\left(\frac{pos}{10000^{2i/d_{\text{model}}}}\right)\]

You aren’t just tagging a token with a index number. When the attention matrix calculates the Query-Key inner product, the mathematical expansion separates into distinct interactions:

\[\text{Attention Score} \propto \underbrace{W_q e (W_k e)^T}_{\text{Semantic-Semantic}} + \underbrace{W_q e (W_k p)^T + W_q p (W_k e)^T}_{\text{Semantic-Positional}} + \underbrace{W_q p (W_k p)^T}_{\text{Positional-Positional}}\]

The network naturally disentangles what the token means from where it sits in the sequence. The exact same principle applies when giving an agent a temporal signal: structure the observation so the math works with the neural network rather than against it.


The Trap of Reward Engineering

If observation design is about clarity, reward design is about incentives—and agents are master incentive hackers.

It’s easy to write a reward function that seems reasonable on paper, only to watch your agent find a bizarre shortcut that completely misses the point. The golden rule of reward design is simple: reward the model only if it achieves the desired objective through the intended pathway.

A common trap occurs when trying to enforce negative constraints. Suppose you are building a tool-use environment in Blender, and you want to prevent an agent from deleting foundational objects in the scene. The natural reaction is to apply a heavy penalty whenever it deletes an object:

Total Reward = Task Completion Bonus - Constraint Penalty

This creates a critical loophole:

  1. Scenario A: The agent solves the task while breaking the constraint, yielding a high task reward that outweighs the penalty.
  2. Scenario B: The agent follows the rules perfectly, but takes a longer path with a lower net reward.

In Scenario A, the agent learns that paying the “fine” is worth it to get the payout. Instead of trying to balance soft penalties in a scalar reward, hard constraints are often better handled by hard environment limits or strict state resets.

   [ Agent Action ]
          │
          ├── (Constraint Violated?) ──► [ Hard Reset / Episode Terminated ]
          │
          └── (Valid Execution)    ──► [ Calculate Dense Objective Reward ]


Termination vs. Truncation: A Subtle Distinction

When an episode ends, how you report that ending to the underlying RL algorithm changes how the agent learns. There are two distinct ways an episode stops:

  1. Termination: The agent hit a natural ending state—it either successfully completed the workflow or catastrophically failed. Because no future actions exist past this point, the estimated value of the final state ($V(S_{\text{final}})$) must be set strictly to 0.
  2. Truncation: The experimenter stepped in and cut the rollout short (e.g., reaching a maximum step count of 50).

If you zero out the state value during a truncation, you send a false signal to the agent that the state just before the time limit was inherently worthless. Instead, truncated states must bootstrap their estimated future value, signaling to the model: “You weren’t done, but we ran out of time.”


Building Robust Tool Environments

When building actionable environments—especially around complex external software like CAD, code interpreters, or 3D engines—a few practical rules separate toy setups from robust environments:

Building environments for reinforcement learning is messy, iterative work. But as we move toward agents that interact with complex software pipelines and multi-step real-world workflows, the real breakthroughs won’t just come from bigger models—they will come from the environments we design for them to explore.