Topic

benchmark

Repositories (1837)

sv-benchmarks
sv-benchmarks sosy-lab

Collection of Verification Tasks (MOVED, please follow the link)

192
gl-bench
gl-bench munrocket JavaScript

⏳ WebGL performance monitor with CPU/GPU load.

192
bench-node
bench-node RafaelGSS JavaScript

A powerful Node.js benchmark library

192
benchmark-php
benchmark-php vanilla-php PHP

:rocket: A benchmark script for PHP and MySQL (Archived)

191
Winrift
Winrift emylfy PowerShell

Windows 11 post-install pipeline — audit, tweak, harden, customize

190
db-benchmarks
db-benchmarks db-benchmarks PHP

Fair database benchmarks framework and datasets

190
metabench
metabench ldionne CMake

A simple framework for compile-time benchmarks

189
jmeter-grpc-request
jmeter-grpc-request zalopay-oss Java

JMeter gRPC Request load test plugin for gRPC

189
uibench
uibench localvoid JavaScript

UI Benchmark

188
FastEval
FastEval FastEval Python

Fast & more realistic evaluation of chat language models. Includes leaderboard.

188
PersonaMem
PersonaMem bowen-upenn Python

[COLM 2025] Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale

186
UBUNTU20-CIS
UBUNTU20-CIS ansible-lockdown YAML

Ansible CIS Benchmark Compliance Remediation for UBUNTU20

186
D-OPTIMIZER
D-OPTIMIZER AveYo Batchfile

Make Dota 2 fps great again

185
scATAC-benchmarking
scATAC-benchmarking pinellolab Jupyter Notebook

Benchmarking computational single cell ATAC-seq methods

185
Single-Image-Deraining
Single-Image-Deraining panda-lab

Single Image Deraining: A Comprehensive Benchmark Analysis

185
Jax-RS-Performance-Comparison
Jax-RS-Performance-Comparison smallnest Java

:zap: Performance Comparison of Jax-RS implementations and embedded containers

184
jsbenchmark
jsbenchmark jsbenchmark Vue

A straightforward JavaScript benchmarking tool and REPL with support for ES modules and libraries.

183
JSONBench
JSONBench ClickHouse Shell

JSONBench: a Benchmark For Data Analytics On JSON

183
UHGEval
UHGEval IAAR-Shanghai Python

[ACL 2024] User-friendly evaluation framework: Eval Suite & Benchmarks: UHGEval, HaluEval, HalluQA, etc.

182
BinKit
BinKit SoftSec-KAIST Shell

Binary Code Similarity Analysis (BCSA) Benchmark

182
memory-maze
memory-maze jurgisp Python

Evaluating long-term memory of reinforcement learning algorithms

181
qpbenchmark
qpbenchmark qpsolvers Python

Benchmark for quadratic programming solvers available in Python

181
Shot2Story
Shot2Story bytedance Python

A new multi-shot video understanding benchmark Shot2Story with comprehensive video summaries and detailed shot-level captions.

179
xFinder
xFinder IAAR-Shanghai Python

[ICLR 2025] xFinder: Large Language Models as Automated Evaluators for Reliable Evaluation

179
VPR-datasets-downloader
VPR-datasets-downloader gmberton Python

Automatic download VPR datasets in a standard format

178
k8s-security-policies
k8s-security-policies raspbernetes Open Policy Agent

This repository offers a comprehensive library of security policies designed to enhance the security of Kubernetes cluster configurations. The policie...

177
OSWorld-G
OSWorld-G xlang-ai TypeScript

[NeurIPS 2025 Spotlight] Scaling Computer-Use Grounding via UI Decomposition and Synthesis

177
MMTrustEval
MMTrustEval thu-ml Python

A toolbox for benchmarking trustworthiness of multimodal large language models (MultiTrust, NeurIPS 2024 Track Datasets and Benchmarks)

177
json-benchmark
json-benchmark serde-rs C++

nativejson-benchmark in Rust

176
freqbench
freqbench kdrag0n Python

Comprehensive CPU frequency performance/power benchmark

176
AMO-Bench
AMO-Bench meituan-longcat Python

This is the official repo for the paper "AMO-Bench: Large Language Models Still Struggle in High School Math Competitions".

176
ecs
ecs andygeiss Go

Build your own Game-Engine based on the Entity Component System concept in Golang.

176
awesome-agent-evolution
awesome-agent-evolution Shiyao-Huang JavaScript

Open survey and evidence map for AI agent evolution, self-evolving agents, memory, skills, harnesses, benchmarks, and agent-swarm systems.

175
DenseMatchingBenchmark
DenseMatchingBenchmark DeepMotionAIResearch Python

Dense Matching Benchmark

174
space_robotics_bench
space_robotics_bench AndrejOrsula Python

Robot Learning Beyond Earth

174
SWE-CI
SWE-CI SKYLENAGE-AI Python

SWE-CI: Evaluating Agent Capabilities in Maintaining Codebases via Continuous Integration

174
ChemLLMBench
ChemLLMBench ChemFoundationModels Jupyter Notebook

Official Code for What can Large Language Models do in chemistry? A comprehensive benchmark on eight tasks (In NeurIPS 2023)

174
fast-crystal
fast-crystal icyleaf Crystal

💨 Writing Fast Crystal 😍 -- Collect Common Crystal idioms.

173
benchmarks
benchmarks catboost Jupyter Notebook

Comparison tools

172
ollama-benchmark
ollama-benchmark LarHope Python

Ollama based Benchmark with detail I/O token per second. Python with Deepseek R1 example.

172
bsuccinct-rs
bsuccinct-rs beling Rust

Rust libraries and programs focused on succinct data structures

172
MedXpertQA
MedXpertQA TsinghuaC3I Python

[ICML 2025] MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

172
BLINK_Benchmark
BLINK_Benchmark zeyofu Python

This repo contains evaluation code for the paper "BLINK: Multimodal Large Language Models Can See but Not Perceive". https://arxiv.org/abs/2404.12...

171
dlbench
dlbench hclhkbu Python

Benchmarking State-of-the-Art Deep Learning Software Tools

170
ossf-cve-benchmark
ossf-cve-benchmark ossf-cve-benchmark TypeScript

The OpenSSF CVE Benchmark consists of code and metadata for over 200 real life CVEs, as well as tooling to analyze the vulnerable codebases using a va...

170
HPOBench
HPOBench automl Python

Collection of hyperparameter optimization benchmark problems

170
Hardening-Audit-Tool-AuditTAP
Hardening-Audit-Tool-AuditTAP fbprogmbh PowerShell

FBPro Audit Test Automation Package allows you to create compliance reports for your systems. The resulting HTML-reports provide a transparent overvie...

170
face-occlusion-generation
face-occlusion-generation kennyvoo Python

[CVPRW 2022] Delving into High-Quality Synthetic Face Occlusion Segmentation Datasets

170
active_genie
active_genie Roriz Ruby

The Lodash for GenAI: Real Value + Consistent + Model-Agnostic

169
http-router
http-router sunrise-php PHP

A powerful solution as the foundation of your project.

168