Benchmark your local LLMs.
CausalGym: Benchmarking causal interpretability methods on linguistic tasks
PhyX: Does Your Model Have the "Wits" for Physical Reasoning?
Benchmark for voice cloning robustness, speaker privacy, and audio protection across 26 TTS and VC models.
Python .pyc decompiler (3.0–3.14) with a contamination-aware benchmark harness. Rule-only pass + one Codex call per module; evaluated on fuzz-syntheti...
Benchmark.js results in ASCII tables for NodeJS
Measuring the performance of popular streaming engines with Yahoo's Streaming Benchmark
An inquiry into nondogmatic software development. An experiment showing double performance of the code running on JVM comparing to equivalent native C...
Evaluation of Line Detection and Association
Official implementation of PRUDEX-Compass
😌 Find the perfect frontend framework, based on what matters most for your project
CockroachDB examples using Docker and Docker Compose
In this paper, a new stochastic optimizer, which is called slime mould algorithm (SMA), is proposed based upon the oscillation mode of slime mould in...
Web Fuzzing Dataset (WFD): a set of web/enterprise applications for experimentation in automated system testing
An ecosystem for digital reticular chemistry
Benchmark results repository service
This is the issue repository for a typescript framework meant to performance test anything even remotely rest-like and related tools
:hugs: AeroPath: An airway segmentation benchmark dataset with challenging pathology
WFCommons: A Framework for Enabling Scientific Workflow Research and Development
OneEval: Open EvalScope evaluation artifacts for LLMs — subset breakdowns, pass@k curves, and reproducible evaluation protocols.
Benchmarking tool to stress real-time protocols
WebAssembly real-world performance benchmark — iswebassemblyfastyet.com
Hyperspectral and soil-moisture data from a field campaign based on a soil sample. Karlsruhe (Germany), 2017.
Benchmark for evaluating open-ended generation
Aix-bench, the Java benchmark for code synthesis problem.
🔥Performance Wars Benchmarking C# - This repository contains a collection of C# benchmarks to compare the performance of different approaches to solv...
A multi-modal Python library for benchmarking lakehouse engines and ELT scenarios, supporting both industry-standard and novel benchmarks.
Simple non-academic performance comparison of available open source implementations of R-tree spatial index using linear, quadratic and R* balancing a...
Code accompanying our IARAI Weather4cast Challenge
web application that charts and compares multiple frame time logs at the same time. Compatible with FPS benchmarking programs such as PresentMon, OCAT...
Benchmarking various Python package managers
[NeurIPS 2025] Official implementation for the paper "SeePhys: Does Seeing Help Thinking? -- Benchmarking Vision-Based Physics Reasoning"
Apache Benchmark Docker image
SQL-ProcBench is an open benchmark for procedural workloads in RDBMSs.
Unity project showcasing A* pathfinding, fully jobified & burst compiled. It also contains examples of RaycastCommand and BoxcastCommand that are used...
SQLite Benchmark
Federated Learning Framework Benchmark (UniFed)
CSS in JS Benchmarks for React Native
Scripts to run and benchmark scRNA-seq cell cluster labeling methods
Benchmark to compare async web server + interpreter + web client implementations across various languages
Testing out a Zero Cost Abstraction in Rust compared to similar approaches in C# and Java
SQL scripts for HeatWave benchmarking
This package provides a python toolkit for the evaluation on the "SeasonDepth: Cross-Season Monocular Depth Prediction Dataset and Benchmark under Mul...
:zap: Optimizing Python code by implementing a C++ extension
[TBench 2024] Official implementation of "AIGCBench: Comprehensive Evaluation of Image-to-Video Content Generated by AI"
A collection of RISC-V Vector (RVV) benchmarks to help developers write portably performant RVV code. (Results)
[ICCV2025] Extrapolated Urban View Synthesis Benchmark
All-in-One Safety Evaluation Framwork
🔮 Benchmarking and visualization toolkit for penalized Cox models