OpenEnv documentation
REPL Environment for OpenEnv
REPL Environment for OpenEnv
repl_env is a Python REPL environment for Recursive Language Model (RLM) style execution. The model writes code that runs in a persistent namespace and can:
- inspect
context - execute Python across multiple turns with persistent state
- call
llm_query(...)andllm_query_batched(...)to query a language model - call
rlm_query(...)andrlm_query_batched(...)for recursive child runs, when configured - finish with
FINAL(...),FINAL_VAR(...), oranswer = {"content": ..., "ready": True}
The package provides:
REPLEnv: the async client for a remote server (.sync()for synchronous code), withexecute(code),submit_final_answer(answer),get_variable(name)andlist_variables()on top ofreset()/step()LocalREPLEnv: the same environment, in processLocalRLMRunner: a local RLM loop that prompts a model, runs its code and handles recursion
Quick Start
Start a server:
PYTHONPATH=src:envs uvicorn envs.repl_env.server.app:app --host 127.0.0.1 --port 8000
Async:
import asyncio
from repl_env import REPLEnv
async def main():
async with REPLEnv(base_url="http://127.0.0.1:8000") as env:
result = await env.reset(
context="alpha beta gamma",
task_prompt="Count the words",
)
result = await env.execute("count = len(context.split())")
result = await env.execute("print(FINAL(count))")
print(result.done)
asyncio.run(main())Sync:
from repl_env import REPLEnv
with REPLEnv(base_url="http://127.0.0.1:8000").sync() as env:
result = env.reset(
context="alpha beta gamma",
task_prompt="Count the words",
)
result = env.execute("count = len(context.split())")
result = env.execute("print(FINAL(count))")
print(result.observation.result.stdout)In Process
from repl_env import LocalREPLEnv
with LocalREPLEnv() as env:
result = env.reset(
context="The quick brown fox jumps over the lazy dog",
task_prompt="Count the words",
)
result = env.execute("count = len(context.split())")
result = env.execute("print(FINAL(count))")
print(env.state().final_answer)Server Configuration
Environment variables read by server/app.py:
| Variable | Default | Description |
|---|---|---|
HF_TOKEN | unset | Enables llm_query on the server. Without it, a client can pass hf_token to reset() |
LLM_MODEL | Qwen/Qwen3.5-9B | Default model for llm_query. A client can pass llm_model to reset() |
REPL_MAX_ITERATIONS | 30 | Maximum steps per episode |
REPL_MAX_OUTPUT_LENGTH | 8192 | Maximum captured output per step |
REPL_CONTEXT_PREVIEW_LENGTH | 500 | Length of context_preview in observations |
REPL_RLM_MAX_DEPTH | 2 | Maximum recursion depth for rlm_query |
REPL_RLM_MAX_ITERATIONS | 30 | Maximum iterations for recursive child runs |
MAX_CONCURRENT_ENVS | 8 | Maximum concurrent sessions |
Reward
Rewards use the OpenEnv rubric system. The default REPLRubric combines:
- Outcome reward (on terminal steps): compares
final_answeragainstexpected_answerif provided. Returns 1.0 for match, 0.0 otherwise. - Process reward (on non-terminal steps): returns -0.05 for code execution errors, 0.0 for successful steps.
- Failure reward: returns -0.1 when max iterations exhausted without an answer.
For RL training (GRPO, etc.), pass expected_answer at reset time:
with LocalREPLEnv() as env:
env.reset(
context="...",
task_prompt="...",
expected_answer="42", # ground truth for rubric scoring
)
result = env.execute("print(FINAL(42))")
print(result.reward) # 1.0 (correct)Other rubrics: ExactMatchRubric (binary match), FuzzyMatchRubric (1.0 for an exact match, 0.5 when the expected answer is contained in the final answer), CustomMetricRubric (your metric(expected, predicted) -> float) and CodeExecutionRubric (per-step error penalty). Pass one at construction:
from repl_env import LocalREPLEnv, CustomMetricRubric, REPLRubric
def my_metric(expected, predicted):
return 1.0 if expected.strip() == predicted.strip() else 0.0
env = LocalREPLEnv(rubric=REPLRubric(outcome=CustomMetricRubric(my_metric)))Running an RLM Locally
LocalRLMRunner takes any chat_fn(messages, model=None) -> str. It works
with HF Inference API, vLLM, SGLang, Ollama, or any OpenAI-compatible server.
With HF Inference API:
from huggingface_hub import InferenceClient
from repl_env import LocalRLMRunner, RLM_SYSTEM_PROMPT
client = InferenceClient(model="Qwen/Qwen3.5-9B", timeout=300)
def chat_fn(messages, model=None):
response = client.chat.completions.create(
model=model or "Qwen/Qwen3.5-9B",
messages=messages,
max_tokens=2048,
temperature=0.6,
extra_body={"chat_template_kwargs": {"enable_thinking": False}},
)
return response.choices[0].message.content
runner = LocalRLMRunner(chat_fn, max_iterations=30, max_depth=2)
result = runner.run("The answer is 42", "What number is mentioned?")
print(result.final_answer)With a local vLLM server:
from openai import OpenAI
from repl_env import LocalRLMRunner
client = OpenAI(base_url="http://localhost:8000/v1", api_key="unused")
def chat_fn(messages, model=None):
response = client.chat.completions.create(
model=model or "Qwen/Qwen3.5-9B",
messages=messages,
max_tokens=2048,
temperature=0.6,
)
return response.choices[0].message.content
runner = LocalRLMRunner(chat_fn, max_iterations=30, max_depth=2)
result = runner.run(context, task)Different Models for Outer and Inner Loops
The outer loop (code generation) can use a large model while inner llm_query/rlm_query calls use a smaller, faster model. Pass a
custom backend_factory to the runner:
from openai import OpenAI
from huggingface_hub import InferenceClient
from repl_env import LocalRLMRunner
from repl_env.recursive_backends import BackendLimits, LocalChildRLMBackend
# Outer loop: large local model via vLLM
vllm = OpenAI(base_url="http://localhost:8000/v1", api_key="unused")
def outer_chat(messages, model=None):
r = vllm.chat.completions.create(
model="Qwen/Qwen3-32B", messages=messages, max_tokens=2048,
)
return r.choices[0].message.content
# Inner calls (llm_query/rlm_query): smaller HF-hosted model
hf = InferenceClient(model="Qwen/Qwen3.5-9B")
def inner_chat(messages, model=None):
r = hf.chat.completions.create(
model=model or "Qwen/Qwen3.5-9B", messages=messages, max_tokens=2048,
extra_body={"chat_template_kwargs": {"enable_thinking": False}},
)
return r.choices[0].message.content
def my_backend_factory(llm_chat_fn, **kwargs):
return LocalChildRLMBackend(
inner_chat, # inner calls use the smaller model
runner_factory=LocalRLMRunner,
system_prompt=kwargs["system_prompt"],
max_iterations=kwargs["max_iterations"],
env_max_iterations_multiplier=kwargs["env_max_iterations_multiplier"],
depth=kwargs["depth"],
limits=BackendLimits(max_depth=2),
)
runner = LocalRLMRunner(
outer_chat, # outer loop: large model
backend_factory=my_backend_factory, # inner calls: small model
max_iterations=30,
max_depth=2,
)
result = runner.run(context, task)LocalRLMRunner also takes recursion limits (max_children_total, max_children_per_batch, per_child_timeout_s, result_truncation_limit) and lifecycle callbacks (on_subcall_start(depth, model, prompt_preview), on_subcall_complete(depth, model, duration, error_or_none)). Its results carry lightweight child trace metadata.
Actions and Observations
REPLAction
code: str = ""
is_final: bool = False
final_answer: str | None = NoneREPLObservation
result: CodeBlockResult
context_preview: str | None
context_length: int
available_variables: list[str]
iteration: int
max_iterations: int
done: bool
reward: float | None
metadata: dictREPL Helpers
When configured, the REPL namespace exposes:
llm_query(prompt, model=None)andllm_query_batched(prompts, model=None)rlm_query(prompt, model=None)andrlm_query_batched(prompts, model=None): recursive child runs. At the maximum depth they fall back to direct model calls.FINAL(value),FINAL_VAR(name)andSHOW_VARS()
Finalization Patterns
FINAL(...)
result = env.execute("answer = 42")
result = env.execute("print(FINAL(answer))")FINAL_VAR(...)
result = env.execute("my_answer = '42'")
result = env.execute('print(FINAL_VAR("my_answer"))')answer dict
result = env.execute("answer['content'] = '42'")
result = env.execute("answer['ready'] = True")Prompts and Examples
prompts.py has the system prompts and helpers used by the runner: RLM_SYSTEM_PROMPT, RLM_SYSTEM_PROMPT_QWEN, QueryMetadata, build_rlm_system_prompt(...), build_user_prompt(...), extract_code_blocks(...) and format_observations(...).
Examples, which default to Qwen/Qwen3.5-9B through Hugging Face inference (needs HF_TOKEN):