reinforcement-learning
llama3
What this is
8B parameters. A policy trained by trial and reward rather than on a text corpus.
llama3 — the lab's own terms rather than a standard licence. Worth reading before you ship.
Release chainonly the steps this archive can prove
- weights posted2025-04-02
Measured herefrom our own daily capture — upstream publishes today only · one day in this window was never captured
← the board