Topic

benchmark

Repositories (1891)

php-di-container-benchmarks
php-di-container-benchmarks kocsismate PHP

Benchmark for some popular PHP Dependency Injection Containers.

77
python-pytest-harvest
python-pytest-harvest smarie Python

Store data created during your `pytest` tests execution, and retrieve it at the end of the session, e.g. for applicative benchmarking purposes.

77
LruClockCache
LruClockCache tugrul512bit C++

A low-latency LRU approximation cache in C++ using CLOCK second-chance algorithm. Multi level cache too. Up to 2.5 billion lookups per second.

77
rsb
rsb gamelife1314 Rust

a http server benchmark tool written in rust 🦀

77
AR-Touch
AR-Touch erfansn Kotlin

🔮 Obtain the power of touchless interaction with display screens

77
DocumentFileCompat
DocumentFileCompat ItzNotABug Kotlin

Blazing fast AndroidX DocumentFile alternative for Android SAF (scoped storage). Up to ~14x faster on large directories.

77
Okutama-Action
Okutama-Action miquelmarti CSS

Okutama-Action: An Aerial View Video Dataset for Concurrent Human Action Detection

76
ipc_benchmark
ipc_benchmark detailyang Python

IPC benchmark on Linux

76
go-cache-benchmark
go-cache-benchmark vmihailenco Go

Cache benchmark for Golang

76
MEDFAIR
MEDFAIR ys-zong Python

[ICLR 2023 spotlight] MEDFAIR: Benchmarking Fairness for Medical Imaging

76
llm-benchmark
llm-benchmark terryyz

A list of LLM benchmark frameworks.

76
CreativeBench
CreativeBench ZethWang Python

Official code for "CreativeBench: Benchmarking and Enhancing Machine Creativity via Self-Evolving Challenges"

76
stepshield
stepshield glo26 Python

StepShield: When, Not Whether to Intervene on Rogue Agents — NeurIPS 2026 benchmark for temporal evaluation of AI agent guardrails (9,429 trajectories...

76
MedMemoryBench
MedMemoryBench AQ-MedAI Python

The source code and data link of paper "MedMemoryBench: Benchmarking Agent Memory in Personalized Healthcare".

76
aa-agentperf-local
aa-agentperf-local ArtificialAnalysis Python

Benchmark local LLM serving by replaying real agent trajectories

76
benchable
benchable MatheusRich Ruby

Write benchmarks without the hassle.

75
TLCBench
TLCBench tlc-pack Python

Benchmark scripts for TVM

75
RePlan
RePlan JIA-Lab-research Python

(ECCV2026) RePlan: Reasoning-Guided Region Planning for Complex Instruction-Based Image Editing

75
the-cpp-abstraction-penalty
the-cpp-abstraction-penalty germandiagogomez C++

Modern C++ benchmarking

74
rpc-bench
rpc-bench bp-alex Java

RPC Benchmark of gRPC, Aeron and KryoNet

74
BenchmarkFcns
BenchmarkFcns mazhar-ansari-ardeh C++

A Python and MATLAB implementation of mathematical test functions for benchmarking optimization algorithms.

74
Turbo-Histogram
Turbo-Histogram powturbo C

Fastest Histogram Construction

74
KoMT-Bench
KoMT-Bench LG-AI-EXAONE Python

Official repository for KoMT-Bench built by LG AI Research

74
vibe
vibe vector-index-bench Python

Vector Index Benchmark for Embeddings (VIBE) is an extensible benchmark for approximate nearest neighbor search methods, or vector indexes, using mod...

74
pdf-text-extraction-benchmark
pdf-text-extraction-benchmark ckorzen TeX

A project about benchmarking and evaluating existing PDF extraction tools on their semantic abilities to extract the body texts from PDF documents, es...

73
CMI
CMI zju-vipa Python

[IJCAI-2021] Contrastive Model Inversion for Data-Free Knowledge Distillation

73
generalization
generalization lechmazur

Thematic Generalization Benchmark: measures how effectively various LLMs can infer a narrow or specific "theme" (category/rule) from a small set of ex...

73
conflictbank
conflictbank zhaochen0110 Python

Code and data for "ConflictBank: A Benchmark for Evaluating the Influence of Knowledge Conflicts in LLM" (NeurIPS 2024 Track Datasets and Benchmarks)

73
rust-storage-bench
rust-storage-bench marvin-j97 Rust

Benchmarking Rust key-value storage engines

73
web-components-benchmark
web-components-benchmark vogloblinsky JavaScript

Web Components benchmark for a various Web Components technologies

72
ncnn-benchmark
ncnn-benchmark BUG1989 CMake

The benchmark of ncnn that is a high-performance neural network inference framework optimized for the mobile platform

72
scalajs-benchmark
scalajs-benchmark japgolly Scala

Benchmarks: write in Scala or JS, run in your browser. Live demo:

72
icor-codon-optimization
icor-codon-optimization Lattice-Automation Python

RNN-based Codon Optimization Tool. Publication: https://doi.org/10.1186/s12859-023-05246-8

72
One-shot-Human-Parsing
One-shot-Human-Parsing Charleshhy Python

[AAAI 2021] (oral) Progressive One-shot Human Parsing, [TPAMI 2023] End-to-end One-shot Human Parsing

72
GenExam
GenExam OpenGVLab Python

[ICML 2026] GenExam: A Multidisciplinary Text-to-Image Exam

72
o11y-bench
o11y-bench grafana Python

Open benchmark on observability tasks built on Harbor

72
MAVBench
MAVBench harvard-edge Python

Simulator + benchmark suite for Micro Aerial Vehicle design.

71
ChemBench
ChemBench shenwanxiang HTML

MoleculeNet benchmark dataset & MolMapNet dataset

71
go-interface-values
go-interface-values akutz Go

When storing a value in a Go interface allocates memory on the heap.

71
TaskMeAnything
TaskMeAnything JieyuZ2 Python

[NeurIPS 2024] A task generation and model evaluation system for multimodal language models.

71
ESBench
ESBench ESBenchmark TypeScript

Modern JavaScript benchmarking tool.

71
All-Angles-Bench
All-Angles-Bench Chenyu-Wang567 Python

Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs

71
Image-Compression-Benchmark
Image-Compression-Benchmark WangXuan95 Python

A comparison of many lossless image compression formats.

71
ageval
ageval ZJU-REAL Python

Agent eval on one running base. Swap the agent under test with plugins; run the same dataset anywhere.

71
SRsurvey
SRsurvey saeed-anwar

A Deep Journey into Super-resolution: A Survey, ACM Computing Surveys

70
quic_vs_tcp
quic_vs_tcp Shenggan Python

A Survey and Benchmark of QUIC

70
untangle
untangle bmucsanyi Python

Large-scale uncertainty benchmark in deep learning.

70
CourtSI
CourtSI Visionary-Laboratory Python

Stepping VLMs onto the Court: Benchmarking Spatial Intelligence in Sports

70
Audio-Oscar
Audio-Oscar ziye26 Python

Audio-Oscar is a multi-agent framework for generating long-form, controllable audio from complex audio scene descriptions.

70
apple-silicon-llm-bench
apple-silicon-llm-bench john-rocky Python

Reproducible on-device LLM benchmarks for Apple Silicon (iPhone 17 Pro, M4 Max): Apple Core AI, MLX, llama.cpp, LiteRT-LM and Core ML on the same mode...

70