Benchmarking repository-borne prompt injection attacks and lightweight defenses for local coding agents. DL4C @ ICML 2026.
Java Hashing, CRC and Checksum Benchmark (JMH)
Tyler's Frame Machine is a simple, free, educational, and portable tool for testing, benchmarking, comparison, and demonstration. TFM supports OpenGL,...
Golang logging library benchmarks
Simple benchmark for testing your DOM diffing algorithm.
Built a smart beta portfolio and compared it to a benchmark index by calculating the tracking error. Built a portfolio using quadratic programming to...
The official code of EMNLP 2022, "SCROLLS: Standardized CompaRison Over Long Language Sequences".
⚡️📊 Compare the performance of Rust project branches
Python project template with unit-tests, documentation, ci-testing and workflows.
[ICLR'25] DataGen: Unified Synthetic Dataset Generation via Large Language Models
ToMBench: Benchmarking Theory of Mind in Large Language Models, ACL 2024.
This is the official repo of MLLM-CL.
The official repo of "MACRO: Advancing Multi-Reference Image Generation with Structured Long-Context Data"
ChineseHarm-Bench: A Chinese Harmful Content Detection Benchmark
Benchmarks for crypto libraries (in Rust, or with Rust bindings)
Benchmarks for common embedded Java and Kotlin web frameworks
Improved docker Golang module dependency cache for faster builds.
This repository is the main Food Recognition Benchmark template and Starter kit. Clone the repository to compete now!
Dataset and benchmark for assessing LLMs in translating natural language descriptions of planning problems into PDDL
A package for benchmarking synthetic relational data generation methods
DafnyBench: A Benchmark for Formal Software Verification
ImputeGAP is a comprehensive Python library for imputation of missing values in time series data. It implements user-friendly APIs to easily visualize...
ATM-Bench: A benchmark for long-term personalized memory QA spanning ~4 years of multimodal data (images, videos, emails). Features referential querie...
Open-source benchmark datasets and pretrained transformer models in the Filipino language.
A LLM training and evaluation benchmark for credit scoring
One-shot LLM eval cases by NAGI STUDIO - same prompt, different agents (model + harness), runnable artifacts side by side.
A Benchmark for Evaluating Agent Self-Evolution on Real Business Tasks
Your benchmark assistant, written in Go.
Benchmarks for intrinsic word embeddings evaluation.
The NAS Parallel Benchmarks for evaluating C++ parallel programming frameworks on shared-memory architectures
Official PHP benchmark suite
Safe Multi-Agent Isaac Gym benchmark for safe multi-agent reinforcement learning research.
Vision Benchmark in Rome Development Kit
Benchmarking Semantic Query Processing Engines
[NeurIPS 2025] MedAgentBoard: Benchmarking Multi-Agent Collaboration with Conventional Methods for Diverse Medical Tasks
Χ-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?
Self-evolving agentic reward framework for image-editing evaluation — 47.4% on EditReward-Bench from only 100 preference demos, no reward-model traini...
A Redis dehydrator module
The ultimate bandwidth benchmark
Code base for the SENSORIUM competition.
Datasets for Evaluation on Domain Knowledge Graph
A toolbox for benchmarking SOTA discriminative and generative geometry estimation models.
Cookbook to provide solutions to common tasks and problems in using Polars with R
Unity Netcode/Network Benchmark Comparison. Fusion, Fishnet, Mirror, Mirage, Netick, NGO
GeoGuessr benchmark for language models
Spatiotemporal datasets collected for network science, deep learning and general machine learning research.
Micro-benchmarking library for C and C++ with PMU counters tracking
Code and models for the WACV 2021 paper "Benchmark for evaluating pedestrian action prediction"
Mutual information estimators and benchmark