When you first dive into reinforcement learning, it is tempting to think that the magic lies almost entirely in the algorithm. We obsess over policy gradients, reward scaling, and model architecture. But the more you work on training agents to operate in complex domains—whether that means navigating a codebase or sculpting a scene in Blender—the more you realize a stark truth: an agent’s capability is fundamentally bounded by the world you build for it.
Recently, as Large Language Models have started pushing past the static limits of supervised pre-training, there’s been a massive revival in RL environments. People want models that can explore outside their training distribution, figure out novel pathways, and use tools iteratively. But building a proper “RL harness” is an art form—one that sits right at the intersection of system design, math, and subtle software engineering.
Here are a few core insights and lessons learned on how to approach environment design, why simple intuition often breaks down, and what it actually takes to train agents that learn.
In textbook RL, we often pretend the agent looks out at the world and sees everything exactly as it is. In practice, real environments almost never work this way.
An agent doesn’t receive the environment’s true internal state; it receives an observation—a partial, noisy, and often redundant glimpse of reality. If you give an agent a single snapshot of a scene, it loses all concept of velocity, acceleration, and intent. To make intelligent decisions, the agent needs a sequence of observations over time to infer the underlying trajectory.
┌────────────────────────┐
│ Environment State │ (Hidden / Complete)
└───────────┬────────────┘
│
[ Filter / Mask ]
│
▼
┌──────────────────────────────────────┐
│ Observation Stream: (o_t-2, o_t-1) │ ──► Agent Decision
└──────────────────────────────────────┘
This applies to time bounds as well. If an agent has a deadline, it needs to know how much time is left. A common pattern in environment design is encoding cyclical variables—like remaining execution windows or continuous physical loops—into smooth sine and cosine embeddings.
This idea of passing cyclical time to an agent mirrors how Transformers handle sequence order. When you add fixed positional encodings to token embeddings:
\[PE_{(pos, 2i)} = \sin\left(\frac{pos}{10000^{2i/d_{\text{model}}}}\right), \quad PE_{(pos, 2i+1)} = \cos\left(\frac{pos}{10000^{2i/d_{\text{model}}}}\right)\]You aren’t just tagging a token with a index number. When the attention matrix calculates the Query-Key inner product, the mathematical expansion separates into distinct interactions:
\[\text{Attention Score} \propto \underbrace{W_q e (W_k e)^T}_{\text{Semantic-Semantic}} + \underbrace{W_q e (W_k p)^T + W_q p (W_k e)^T}_{\text{Semantic-Positional}} + \underbrace{W_q p (W_k p)^T}_{\text{Positional-Positional}}\]The network naturally disentangles what the token means from where it sits in the sequence. The exact same principle applies when giving an agent a temporal signal: structure the observation so the math works with the neural network rather than against it.
If observation design is about clarity, reward design is about incentives—and agents are master incentive hackers.
It’s easy to write a reward function that seems reasonable on paper, only to watch your agent find a bizarre shortcut that completely misses the point. The golden rule of reward design is simple: reward the model only if it achieves the desired objective through the intended pathway.
A common trap occurs when trying to enforce negative constraints. Suppose you are building a tool-use environment in Blender, and you want to prevent an agent from deleting foundational objects in the scene. The natural reaction is to apply a heavy penalty whenever it deletes an object:
Total Reward = Task Completion Bonus - Constraint Penalty
This creates a critical loophole:
In Scenario A, the agent learns that paying the “fine” is worth it to get the payout. Instead of trying to balance soft penalties in a scalar reward, hard constraints are often better handled by hard environment limits or strict state resets.
[ Agent Action ]
│
├── (Constraint Violated?) ──► [ Hard Reset / Episode Terminated ]
│
└── (Valid Execution) ──► [ Calculate Dense Objective Reward ]
When an episode ends, how you report that ending to the underlying RL algorithm changes how the agent learns. There are two distinct ways an episode stops:
If you zero out the state value during a truncation, you send a false signal to the agent that the state just before the time limit was inherently worthless. Instead, truncated states must bootstrap their estimated future value, signaling to the model: “You weren’t done, but we ran out of time.”
When building actionable environments—especially around complex external software like CAD, code interpreters, or 3D engines—a few practical rules separate toy setups from robust environments:
Building environments for reinforcement learning is messy, iterative work. But as we move toward agents that interact with complex software pipelines and multi-step real-world workflows, the real breakthroughs won’t just come from bigger models—they will come from the environments we design for them to explore.