Topic

benchmark

Repositories (1763)

coir
coir CoIR-team Python

(ACL 2025 Main) A Comprehensive Benchmark for Code Information Retrieval.

151
plf_nanotimer
plf_nanotimer mattreecebentley C++

A simple C++ 03/11/etc timer class for ~microsecond-precision cross-platform benchmarking. The implementation is as limited and as simple as possible...

151
NAS-Benchmark
NAS-Benchmark antoyang Python

[ICLR 2020] NAS evaluation is frustratingly hard

150
EasyIterator
EasyIterator TheLartians C++

🏃 Iterators made easy! Zero cost abstractions for designing and using C++ iterators.

150
pddl-instances
pddl-instances potassco Common Lisp

🌍 PDDL instances covering the International Planning Competitions

150
serverless-faas-workbench
serverless-faas-workbench ddps-lab Python

FunctionBench : A Suite of Workloads for Serverless Cloud Function Service

148
Windows-2019-CIS
Windows-2019-CIS ansible-lockdown YAML

Automated CIS Benchmark Compliance Remediation for Windows Server 2019 with Ansible

148
rvv-bench
rvv-bench camel-cdr Assembly

A collection of RISC-V Vector (RVV) benchmarks to help developers write portably performant RVV code

148
mqperf
mqperf softwaremill Scala
147
bucketbench
bucketbench estesp Go

Go-based framework for running benchmarks against Docker, containerd, runc, or any CRI-compliant runtime

147
xVerify
xVerify IAAR-Shanghai Jupyter Notebook

xVerify: Efficient Answer Verifier for Reasoning Model Evaluations

147
aurora
aurora wenhaochai Python

[ICLR 2025] AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

147
CharXiv
CharXiv princeton-nlp Python

[NeurIPS 2024] CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs

147
aacr-bench
aacr-bench alibaba Python

An Alibaba open-source multi-language benchmark for evaluating LLMs in repository-level automatic code review, featuring an AI-assisted and expert-ver...

147
ClassEval
ClassEval FudanSELab Python

Benchmark ClassEval for class-level code generation.

146
gameworld
gameworld gameworld-project Python

GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents

146
goku
goku jcaromiq Rust

Goku is an HTTP load testing application written in Rust

146
docile
docile rossumai Python

DocILE: Document Information Localization and Extraction Benchmark

146
golang-benchmarks
golang-benchmarks SimonWaldherr Go

Go(lang) benchmarks - (measure the speed of golang)

145
space_robotics_bench
space_robotics_bench AndrejOrsula Python

Robot Learning Beyond Earth

145
TCPDBench
TCPDBench alan-turing-institute

The Turing Change Point Detection Benchmark: An Extensive Benchmark Evaluation of Change Point Detection Algorithms on real-world data

144
php-orm-benchmark
php-orm-benchmark kenjis PHP

PHP ORM Benchmark

143
benchmarks
benchmarks lmdbjava Shell

Benchmark of open source, embedded, memory-mapped, key-value stores available from Java (JMH)

143
video-quality-metrics
video-quality-metrics CrypticSignal Python

Uses FFmpeg to benchmark video encoders to compare VMAF, SSIM and PSNR with different encoder settings.

143
leaderboard
leaderboard KGQA Jupyter Notebook

You can find the most recent KGQA benchmark numbers from publications here.

143
jsbench-me
jsbench-me psiho

jsbench.me - JavaScript performance benchmarking playground

142
smartbugs-curated
smartbugs-curated smartbugs Solidity

SB Curated is a curated dataset of Solidity smart contracts annotated with tagged vulnerabilities. The dataset was created to evaluate the accuracy of...

142
V2X-Sim
V2X-Sim ai4ce

[RA-L2022] V2X-Sim Dataset and Benchmark

142
SpeedTests
SpeedTests jabbalaci Python

comparing the execution speeds of various programming languages

142
service-mesh-benchmark
service-mesh-benchmark kinvolk Shell
141
ElegantMustard
ElegantMustard lscambo13

An elegant RTSS Overlay to showcase your benchmark stats in style.

139
EmoBench-M
EmoBench-M Emo-gml Python

EmoBench-M: A benchmark for evaluating Emotional Intelligence in Multimodal Large Language Models.

138
Video-Bench
Video-Bench PKU-YuanGroup Python

A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models!

138
facies_classification_benchmark
facies_classification_benchmark yalaudah Python

The repository includes PyTorch code, and the data, to reproduce the results for our paper titled "A Machine Learning Benchmark for Facies Classificat...

138
typescript-orm-benchmark
typescript-orm-benchmark emanuelcasco TypeScript

⚖️ ORM benchmarking for Node.js applications written in TypeScript

137
VCSL
VCSL alipay Python

Video Copy Segment Localization (VCSL) dataset and benchmark [CVPR2022]

137
go-perftuner
go-perftuner go-perf Go

Helper tool for manual Go code optimization.

137
Touchstone
Touchstone MrGiovanni Jupyter Notebook

[NeurIPS 2024] Touchstone - Benchmarking AI on 5,172 o.o.d. CT volumes and 9 anatomical structures

136
chembench
chembench lamalab-org Python

How good are LLMs at chemistry?

136
arewefastyet
arewefastyet mozilla JavaScript

NOT MAINTAINED ANYMORE! New project is located on https://github.com/mozilla-frontend-infra/js-perf-dashboard -- AreWeFastYet is a set of tools used f...

135
PersonaMem
PersonaMem bowen-upenn Python

[COLM 2025] Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale

135
Domain-generalization-fault-diagnosis-benchmark
Domain-generalization-fault-diagnosis-benchmark CHAOZHAO-1 Python

This is a benckmark for domain generalization-based fault diagnosis (基于领域泛化的相关代码)

135
PHP-Frameworks-Bench
PHP-Frameworks-Bench myaaghubi PHP

A library to make benchmarks from PHP frameworks.

135
awesome-world-model-evolution
awesome-world-model-evolution OpenRaiser

A curated collection of research papers, models, and resources tracing the evolution from specialized models to unified world models.

134
xbench-evals
xbench-evals xbench-ai Python

Evergreen, contamination-free, real-world, domain-specific AI evaluation framework

134
actors
actors plokhotnyuk Scala

Evaluation of API and performance of different actor libraries

132
golang-graphql-benchmark
golang-graphql-benchmark appleboy Go

benchmark of golang GraphQL framework.

132
jvm-performance-benchmarks
jvm-performance-benchmarks ionutbalosin Java

Java Virtual Machine (JVM) Performance Benchmarks with a primary focus on top-tier Just-In-Time (JIT) Compilers, such as C2 JIT, Graal JIT, and the Fa...

132
THST
THST tuxalin C++

Templated hierarchical spatial trees designed for high-peformance.

132
contender
contender flashbots Rust

spam EVM execution nodes over JSON-RPC & run benchmarks

132