Domain Adaptation of Base Models + ShadowdarkQA Bench
Investigating the effects of continued pre-training for learning precise mechanical rules of TTRPGs.
Recent content on The Gygax Test
Investigating the effects of continued pre-training for learning precise mechanical rules of TTRPGs.
Can an LLM run a satisfying game of Dungeons & Dragons? The Gygax Test explores whether LLMs can create consistent characters, craft emerging narratives, and run tactical combat across a full campaign.
Welcome to The Gygax Test The Gygax Test is a research blog exploring whether Large Language Models (LLMs) can effectively serve as Dungeon Masters in tabletop roleplaying games like Dungeons & Dragons. The Gygax Test evaluates whether an AI can: Create and maintain consistent characters across a long campaign Develop emerging narrative arcs that respond to player choices Run tactically…
Project Releases This page will contain links to code repositories, datasets, model releases, and evaluation frameworks as they become available during the course of this research. Shadowdark QA Bench One A set of QA for Shadowdark RPG across multiple categories, used to gauge knowledge retention of RPG rules. Upcoming Releases GM-Eval Framework : A set of metrics and evaluation scenarios for…