118B parameters. Writes text from a prompt. This is the kind of model people mean by "an LLM".
Packaged as AWQ. AWQ needs a GPU serving stack such as vLLM.
llama2 — the lab's own terms rather than a standard licence. Worth reading before you ship.