vlmrl

An experimental reinforcement learning gym for vision-language models, written in JAX. It lets you plug in environments and models and train them with PPO.

Share on XLicense: Apache-2.0

Overview

vlm-gym is an experimental reinforcement learning gym for vision-language models, written in JAX. The idea is to drop in any environment and any model and train with PPO. It ships with environments such as GeoGuessr, NLVR2 and captioning, and a reference Qwen3-VL-4B-Instruct model that can be converted from Hugging Face to JAX. The trainer, rollout engine and evaluation harness live in a core folder, and GeoGuessr training uses a staged curriculum.

Key features

  • Pluggable vision environments: GeoGuessr, NLVR2, captioning
  • PPO trainer for vision-language models
  • Rollout engine and evaluation against a Hugging Face baseline
  • Converts Qwen3-VL-4B-Instruct from Hugging Face to JAX

Best for

Suited to researchers experimenting with reinforcement learning on vision-language models in JAX. The README marks its status as experimental.

Upstream
sdan/vlm-gym
Fork on GitHub
Guo-astro/vlmrl
Upstream stars
149
Category
AI agents and LLM tools
Language
Python
License
Apache-2.0
Forked
2025-10-14
Sync status
In syncLast synced 2026-09-29