Narrative Goldmine

Home

❯

working

❯

Evaluation benchmarks and leaderboards

Evaluation benchmarks and leaderboards

01 Oct 20261 min read

Properties

Type
  • Note
Status
  • stable
Generated
  • by: process:vault-migrate/1.0 · at: 2026-09-22T12:41:58.55607396Z

1719268663052.jpeg

  • https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard
  • https://huggingface.co/spaces/CultriX/Alt_LLM_LeaderBoard
  • https://huggingface.co/spaces/opencompass/open_vlm_leaderboard
  • https://tatsu-lab.github.io/alpaca_eval/
  • https://huggingface.co/spaces/bigcode/bigcode-models-leaderboard
  • https://huggingface.co/spaces/NPHardEval/NPHardEval-leaderboard
  • https://paperswithcode.com/sota/code-generation-on-humaneval
  • https://evalplus.github.io/leaderboard.html
  • https://eqbench.com/
  • https://ayumi.m8geil.de/

Graph View

Created with Quartz v4.5.2 © 2026

  • Ontology (Turtle)
  • Search index
  • JSON-LD context
  • Explorer