RSS Amplifier

Blog

VIJIL

Helps organizations build and operate agents that humans can trust

vijil.substack.comSource feed ↗3 posts

Dormant Last read · last published · next check
Read 10 days ago and current, but nothing has been published for 24 months.

Elsewhere

Latest posts

Get your MMLU score 20X cheaper and 1000x faster

When developers want to understand how well different Large Language Models (LLM) perform across a common set of tasks, they turn to standard benchmarks such as Massive Multitask Language Understanding (MMLU) and Grade School Math 8K (GSM8K).

Vijil emerges from stealth

To help organizations build AI agents that humans can trust

Evaluating Meta Llama 3 for Stereotyping

Vijil evaluation of stereotyping in Meta Llama models reveals surprising results