A Chinese National Medical Licensing Examination dataset and large languge model benchmarks
Agentic Benchmark for LLM-Crafted Heuristics in Combinatorial Optimization (ICLR'26)
(EMNLP 2025 Findings) Source Evaluation scripts for Humanity's Last Code Exam
[CVPR 2026 Oral] PAI-Bench: A Comprehensive Benchmark for Physical AI
[NeurIPS 2023] A faithful benchmark for vision-language compositionality
AgenticVBench: Can AI Agents Complete Real-World Post-Production Tasks?
[NAACL 2024] MMC: Advancing Multimodal Chart Understanding with LLM Instruction Tuning
BIRL: Benchmark on Image Registration methods with Landmark validations
:trophy: Delightful Benchmarking & Performance Testing
Latency Benchmarking tool
Unified Multi-modal IAA Baseline and Benchmark
Repo for "Physion: Evaluating Physical Prediction from Vision in Humans and Machines", presented at NeurIPS 2021 (Datasets & Benchmarks track)
DTB70 -- A Drone Tracking Benchmark
🪑 Benchmark different versions of same or similar gems & Static Gemfile and installed gem library source code analysis
Code and models for "Pano3D: A Holistic Benchmark and a Solid Baseline for 360 Depth Estimation", OmniCV Workshop @ CVPR21.
Comparison and benchmark of JavaScript serialization libraries (Protocol Buffer, Avro, BSON, etc.)
This repository contains the implementation for the paper "Revisiting Few Shot Object Detection with Vision-Language Models"
NEXBENCH measures what an autonomous agent can actually do on-chain — execute transactions, route swaps, bridge funds, manage DeFi positions, research...
[ICLR 2025] Official implementation and benchmark evaluation repository of <PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical...
RSI-Exam: Measuring Recursive Self-Improvement on Long-Horizon, Executable Research Tasks · Discord https://discord.gg/FBzhepyE
High fidelity benchmark runner
Benchmarking framework for index structures on persistent memory
Serverreview Benchmark Script v3
Fastest Trie structure (Linux & Windows)
[NeurIPS 2024 D&B] Point Cloud Matters: Rethinking the Impact of Different Observation Spaces on Robot Learning
HTTP Load Generator
Microbenchmarks comparing the Julia Programming language with other languages
Command-line DNS benchmark
SUES-200: A Multi-height Multi-scene Cross-view Image Benchmark Across Drone and Satellite
Code for MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems
Run a 110B on a 2016 PC with 16 GB RAM. Know your tok/s before you download. Placement beats budget: predicts speed + memory fit for any GGUF on your...
[ICLR 2026] Official Implementation of "FeatureBench: Benchmarking Agentic Coding for Complex Feature Development"
[IJCAI 2024] FactCHD: Benchmarking Fact-Conflicting Hallucination Detection
A Consistent Data Visualizer for Agents and Humans
[NeurIPS 2024] Evaluation harness for SWT-Bench, a benchmark for evaluating LLM repository-level test-generation
(NeurIPS 2024) Official PyTorch implementation of LOVA3
It is a collection of php benchmarks
Client Code Examples, Use Cases and Benchmarks for Enterprise h2oGPTe RAG-Based GenAI Platform
[ECCV 2024] Official PyTorch Implementation of "How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs"
A Karma plugin to run Benchmark.js over multiple browsers with CI compatible output.
A Python tool to evaluate the performance of VLM on the medical domain.
UME::SIMD A library for explicit simd vectorization.
BioDSA: Framework for Vibe Prototyping of AI Agents for Biomedicine
Multi-Agent Step Race Benchmark: Assessing LLM Collaboration and Deception Under Pressure. A multi-player “step-race” that challenges LLMs to engage i...
Check your internet speed/bandwidth right from your terminal. Built on Golang using chromedp
Locust4j is a load generator for locust, written in Java.
[IJRR2024] The official repository for the WildScenes: A Benchmark for 2D and 3D Semantic Segmentation in Natural Environments
Benchmarking programming languages and web frameworks.
An agent skill that plays ARC-AGI-3. One rule: say what an action will do before you spend it. Claude Code on Opus 5 finished all 25 public games at 1...
RNN benchmarks of pytorch, tensorflow and theano