Skip to content

Latest commit

 

History

20 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Measuring what Matters: Construct Validity in Large Language Model Benchmarks

Measuring what Matters is a systematic review of construct validity practices in benchmarks for LLMs. This repository contains the code and data from the review.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages