Logs performance benchmark repo: Comparing Elastic, Loki and SigNoz
The benchmark to compare performance of PHP ORM solutions.
List of Ruby Tools for doing Performance.
Libsodium WebAssembly benchmarks results.
Learned Sort: a model-enhanced sorting algorithm
A benchmark framework based on Golang
🚀 A comprehensive performance comparison benchmark between different .NET collections.
Framework for benchmarking fully-managed vector databases
CPU micro benchmarks
Automated Benchmarking System for Vitess
Lakehouse storage system benchmark
[NeurIPS 2024] Terra: A Multimodal Spatio-Temporal Dataset Spanning the Earth
Benchmark evaluation code for "SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal" (ICLR 2025)
Human Benchmark is a Flutter app for Android that features many tests to assess your abilities.
This repo contains the code for "MEGA-Bench Scaling Multimodal Evaluation to over 500 Real-World Tasks" [ICLR 2025]
A tabular data visualization engine from your local to CI/CD pipeline
SUES-200: A Multi-height Multi-scene Cross-view Image Benchmark Across Drone and Satellite
Compare performance of macOS browsers based on Speedometer 3.1
Benchmark for some popular PHP Dependency Injection Containers.
a http server benchmark tool written in rust 🦀
🚀 Spiko is a fast, Rust-based load testing tool with a beautiful TUI for real-time insights.
[ECCV 2024] WiMANS: A Benchmark Dataset for WiFi-based Multi-user Activity Sensing
Enable Comprehensive LLM Evaluation on Graph Reasoning
Data race benchmark suite for evaluating OpenMP correctness tools aimed to detect data races.
Cache benchmark for Golang
Store data created during your `pytest` tests execution, and retrieve it at the end of the session, e.g. for applicative benchmarking purposes.
A low-latency LRU approximation cache in C++ using CLOCK second-chance algorithm. Multi level cache too. Up to 2.5 billion lookups per second.
LeakDB (Leakage Diagnosis Benchmark) is a realistic leakage dataset for water distribution networks. The dataset is comprised of a large number of art...
[ICLR26 Oral] RealPDEBench: A Benchmark for Complex Physical Systems with Paired Real-World and Simulated Data
[NeurIPS 2024] Evaluation harness for SWT-Bench, a benchmark for evaluating LLM repository-level test-generation
Code and data of the EMNLP 2022 paper "Why Should Adversarial Perturbations be Imperceptible? Rethink the Research Paradigm in Adversarial NLP".
Write benchmarks without the hassle.
IPC benchmark on Linux
Blazing fast AndroidX DocumentFile alternative for Android SAF (scoped storage). Up to ~14x faster on large directories.
[IJCAI-2021] Contrastive Model Inversion for Data-Free Knowledge Distillation
Modern C++ benchmarking
Okutama-Action: An Aerial View Video Dataset for Concurrent Human Action Detection
[ICLR 2026] IVEBench - Benchmark for Instruction-Guided Video Editing
Benchmark scripts for TVM
[ICLR 2023 spotlight] MEDFAIR: Benchmarking Fairness for Medical Imaging
A list of LLM benchmark frameworks.
Fastest Histogram Construction
Safe Multi-Agent MuJoCo benchmark for safe multi-agent reinforcement learning research.
RPC Benchmark of gRPC, Aeron and KryoNet
[AAAI 2021] (oral) Progressive One-shot Human Parsing, [TPAMI 2023] End-to-end One-shot Human Parsing
Web Components benchmark for a various Web Components technologies
[NeurIPS 2024] A task generation and model evaluation system for multimodal language models.
Benchmarks: write in Scala or JS, run in your browser. Live demo:
Benchmark your 3DS battery