Topic

benchmark

Repositories (1891)

RepoGuardBench
RepoGuardBench DaoyuanLi2816 Python

Benchmarking repository-borne prompt injection attacks and lightweight defenses for local coding agents. DL4C @ ICML 2026.

70
hash-bench
hash-bench bp-alex Java

Java Hashing, CRC and Checksum Benchmark (JMH)

69
TFM
TFM Tylemagne C#

Tyler's Frame Machine is a simple, free, educational, and portable tool for testing, benchmarking, comparison, and demonstration. TFM supports OpenGL,...

69
logbench
logbench rs Go

Golang logging library benchmarks

69
js-diff-benchmark
js-diff-benchmark luwes JavaScript

Simple benchmark for testing your DOM diffing algorithm.

69
smart-beta-portfolio-optimization
smart-beta-portfolio-optimization sanjeevai HTML

Built a smart beta portfolio and compared it to a benchmark index by calculating the tracking error. Built a portfolio using quadratic programming to...

69
scrolls
scrolls tau-nlp Python

The official code of EMNLP 2022, "SCROLLS: Standardized CompaRison Over Long Language Sequences".

69
criterion-compare-action
criterion-compare-action boa-dev JavaScript

⚡️📊 Compare the performance of Rust project branches

69
PythonProjectTemplate
PythonProjectTemplate franneck94 Python

Python project template with unit-tests, documentation, ci-testing and workflows.

69
DataGen
DataGen HowieHwong Python

[ICLR'25] DataGen: Unified Synthetic Dataset Generation via Large Language Models

69
ToMBench
ToMBench zhchen18 Python

ToMBench: Benchmarking Theory of Mind in Large Language Models, ACL 2024.

69
MLLM-CL
MLLM-CL bjzhb666 Python

This is the official repo of MLLM-CL.

69
Macro
Macro HKU-MMLab Python

The official repo of "MACRO: Advancing Multi-Reference Image Generation with Structured Long-Context Data"

69
ChineseHarm-bench
ChineseHarm-bench zjunlp Python

ChineseHarm-Bench: A Chinese Harmful Content Detection Benchmark

69
crypto-bench
crypto-bench briansmith Rust

Benchmarks for crypto libraries (in Rust, or with Rust bindings)

68
http-benchmarks
http-benchmarks orangy Kotlin

Benchmarks for common embedded Java and Kotlin web frameworks

68
golang-docker-cache
golang-docker-cache montanaflynn Dockerfile

Improved docker Golang module dependency cache for faster builds.

68
benchdiff
benchdiff WillAbides Go
68
food-recognition-benchmark-starter-kit
food-recognition-benchmark-starter-kit AIcrowd Jupyter Notebook

This repository is the main Food Recognition Benchmark template and Starter kit. Clone the repository to compete now!

68
planetarium
planetarium BatsResearch Python

Dataset and benchmark for assessing LLMs in translating natural language descriptions of planning problems into PDDL

68
syntherela
syntherela martinjurkovic Python

A package for benchmarking synthetic relational data generation methods

68
DafnyBench
DafnyBench sun-wendy Dafny

DafnyBench: A Benchmark for Formal Software Verification

68
ImputeGAP
ImputeGAP eXascaleInfolab Jupyter Notebook

ImputeGAP is a comprehensive Python library for imputation of missing values in time series data. It implements user-friendly APIs to easily visualize...

68
ATM-Bench
ATM-Bench JingbiaoMei Python

ATM-Bench: A benchmark for long-term personalized memory QA spanning ~4 years of multimodal data (images, videos, emails). Features referential querie...

68
Filipino-Text-Benchmarks
Filipino-Text-Benchmarks jcblaisecruz02 Python

Open-source benchmark datasets and pretrained transformer models in the Filipino language.

67
CALM
CALM The-FinAI Python

A LLM training and evaluation benchmark for credit scoring

67
nagi-bench
nagi-bench nagi-studio HTML

One-shot LLM eval cases by NAGI STUDIO - same prompt, different agents (model + harness), runnable artifacts side by side.

67
GDPevo
GDPevo Prism-Shadow Python

A Benchmark for Evaluating Agent Self-Evolution on Real Business Tasks

67
ben
ben drish Go

Your benchmark assistant, written in Go.

66
word-benchmarks
word-benchmarks vecto-ai

Benchmarks for intrinsic word embeddings evaluation.

66
NPB-CPP
NPB-CPP GMAP C++

The NAS Parallel Benchmarks for evaluating C++ parallel programming frameworks on shared-memory architectures

66
php-version-benchmarks
php-version-benchmarks kocsismate Shell

Official PHP benchmark suite

66
Safe-Multi-Agent-Isaac-Gym
Safe-Multi-Agent-Isaac-Gym chauncygu Python

Safe Multi-Agent Isaac Gym benchmark for safe multi-agent reinforcement learning research.

66
vbr-devkit
vbr-devkit rvp-group Python

Vision Benchmark in Rome Development Kit

66
SemBench
SemBench SemBench Python

Benchmarking Semantic Query Processing Engines

66
MedAgentBoard
MedAgentBoard yhzhu99 Python

[NeurIPS 2025] MedAgentBoard: Benchmarking Multi-Agent Collaboration with Conventional Methods for Diverse Medical Tasks

66
chi-bench
chi-bench actava-ai Python

Χ-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?

66
RewardHarness
RewardHarness TIGER-AI-Lab Python

Self-evolving agentic reward framework for image-editing evaluation — 47.4% on EditReward-Bench from only 100 preference demos, no reward-model traini...

66
ReDe
ReDe TamarLabs C

A Redis dehydrator module

65
TheBandwidthBenchmark
TheBandwidthBenchmark HPC-Dwarfs C

The ultimate bandwidth benchmark

65
sensorium
sensorium sinzlab Jupyter Notebook

Code base for the SENSORIUM competition.

65
OpenBG
OpenBG OpenBGBenchmark

Datasets for Evaluation on Domain Knowledge Graph

65
GeoBench
GeoBench aim-uofa Python

A toolbox for benchmarking SOTA discriminative and generative geometry estimation models.

65
cookbook-rpolars
cookbook-rpolars ddotta CSS

Cookbook to provide solutions to common tasks and problems in using Polars with R

65
unity-netcode-benchmark
unity-netcode-benchmark StinkySteak C#

Unity Netcode/Network Benchmark Comparison. Fusion, Fishnet, Mirror, Mirage, Netick, NGO

65
geobench
geobench ccmdi Python

GeoGuessr benchmark for language models

65
spatiotemporal_datasets
spatiotemporal_datasets benedekrozemberczki

Spatiotemporal datasets collected for network science, deep learning and general machine learning research.

64
b63
b63 okuvshynov C

Micro-benchmarking library for C and C++ with PMU counters tracking

64
PedestrianActionBenchmark
PedestrianActionBenchmark ykotseruba Python

Code and models for the WACV 2021 paper "Benchmark for evaluating pedestrian action prediction"

64
bmi
bmi cbg-ethz Python

Mutual information estimators and benchmark

64