Topic

benchmark

Repositories (1866)

PaddleFleetX
PaddleFleetX PaddlePaddle Python

飞桨大模型开发套件,提供大语言模型、跨模态大模型、生物计算大模型等领域的全流程开发工具链。

482
benchmark
benchmark ordo-one Swift

Swift benchmark runner with many performance metrics and great CI support

479
ProteinGym
ProteinGym OATML-Markslab HTML

Official repository for the ProteinGym benchmarks

469
automlbenchmark
automlbenchmark openml Python

OpenML AutoML Benchmarking Framework

467
mixbench
mixbench ekondis C++

A GPU benchmark tool for evaluating GPUs and CPUs on mixed operational intensity kernels (CUDA, OpenCL, HIP, SYCL, OpenMP)

464
Awesome-Evaluation-of-Visual-Generation
Awesome-Evaluation-of-Visual-Generation ziqihuangg

A list of works on evaluation of visual generation models, including evaluation metrics, models, and systems

463
VulnerableApp
VulnerableApp SasanLabs Java

OWASP VulnerableApp Project: Break it. Scan it. Reproduce it. Benchmark against it. Improve it.

461
DG-PHM
DG-PHM CHAOZHAO-1

This is a reposotory that includes paper、code and datasets about domain generalization-based fault diagnosis and prognosis. (基于领域泛化的故障诊断和...

461
tf_to_trt_image_classification
tf_to_trt_image_classification NVIDIA-AI-IOT Python

Image classification with NVIDIA TensorRT from TensorFlow models.

460
mcpmark
mcpmark eval-sys Python

MCPMark is a comprehensive, stress-testing MCP benchmark designed to evaluate model and agent capabilities in real-world MCP use.

460
LawBench
LawBench open-compass Python

Benchmarking Legal Knowledge of Large Language Models

452
LayoutFrameworkBenchmark
LayoutFrameworkBenchmark layoutBox Swift

Benchmark the performances of various Swift layout frameworks (autolayout, UIStackView, PinLayout, LayoutKit, FlexLayout, Yoga, ...)

448
Craftax
Craftax MichaelTMatthews Python

(Crafter + NetHack) in JAX. ICML 2024 Spotlight.

447
gymfc
gymfc wil3 Python

A universal flight control tuning framework

442
ai-text-detector
ai-text-detector lynote-ai Python

A cautious, explainable AI-like text risk analyzer for local workflows and coding agents.

441
sympact
sympact simonepri JavaScript

🔥 Stupid Simple CPU/MEM "Profiler" for your JS code.

440
prophiler
prophiler fabfuel PHP

PHP Profiler & Developer Toolbar (built for Phalcon)

439
srs-bench
srs-bench ossrs Go

SB(SRS Bench) is a set of benchmark and regression test tools, for SRS and other media servers, supports HTTP-FLV, RTMP, HLS, WebRTC and GB28181.

438
DynamicMap_Benchmark
DynamicMap_Benchmark KTH-RPL Jupyter Notebook

The First Dynamic Map Removal Benchmark | Included 8 SOTA methods | Continous updating

434
benchmarks
benchmarks pnpm JavaScript

Benchmarks of JavaScript Package Managers

433
Awesome_Imputation
Awesome_Imputation WenjieDu Python

Awesome Deep Learning for Time-Series Imputation, including an unmissable paper and tool list about applying neural networks to impute incomplete time...

426
BlurTestAndroid
BlurTestAndroid patrickfav Java

This is a simple App to test some blur algorithms on their visual quality and performance.

423
ChineseBLUE
ChineseBLUE alibaba-research Python

Chinese Biomedical Language Understanding Evaluation benchmark (ChineseBLUE)

423
devtools
devtools crabnebula-dev TypeScript

Inspect and Debug your Tauri applications in style 💃

423
gym-electric-motor
gym-electric-motor upb-lea Python

Gym Electric Motor (GEM): An OpenAI Gym Environment for Electric Motors

423
ronin
ronin Sachini Python

RoNIN: Robust Neural Inertial Navigation in the Wild

423
FedScale
FedScale SymbioticLab Python

FedScale is a scalable and extensible open-source federated learning (FL) platform.

421
pglib-opf
pglib-opf power-grid-lib MATLAB

Benchmarks for the Optimal Power Flow Problem

420
VibeSearchBench
VibeSearchBench VibeBench Python

🔍 The hardest search benchmark in the wild — vague, multi-turn, proactive. 200 long-horizon tasks with persona-driven progressive disclosure, scored...

419
oltpbench
oltpbench oltpbenchmark Java

Database Benchmarking Framework

416
are-we-fast-yet
are-we-fast-yet smarr Java

Are We Fast Yet? Comparing Language Implementations with Objects, Closures, and Arrays

416
SkillEvaluator
SkillEvaluator NVIDIA Python

Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live...

416
pyaf
pyaf antoinecarme Python

PyAF is an Open Source Python library for Automatic Time Series Forecasting built on top of popular pydata modules.

414
OpenRCA
OpenRCA microsoft Python

[ICLR'25] OpenRCA: Can Large Language Models Locate the Root Cause of Software Failures?

414
superpixel-benchmark
superpixel-benchmark davidstutz C++

An extensive evaluation and comparison of 28 state-of-the-art superpixel algorithms on 5 datasets.

414
CRUD_RAG
CRUD_RAG IAAR-Shanghai Python

CRUD-RAG: A Comprehensive Chinese Benchmark for Retrieval-Augmented Generation of Large Language Models

410
modclean
modclean ModClean JavaScript

Remove unwanted files and directories from your node_modules folder

409
vtebench
vtebench alacritty Rust

Generate benchmarks for terminal emulators

408
ros2-performance
ros2-performance irobot-ros C++

Framework to evaluate peformance of ROS 2

408
SKAB
SKAB waico Jupyter Notebook

SKAB - Skoltech Anomaly Benchmark. Time-series data for evaluating Anomaly Detection algorithms.

407
jetson_benchmarks
jetson_benchmarks NVIDIA-AI-IOT Python

Jetson Benchmark

402
gapbs
gapbs sbeamer C++

GAP Benchmark Suite

395
ds4-on-spark
ds4-on-spark Entrpi Shell

Entrpi/ds4, a Blackwell CUDA perf fork of antirez/ds4 on NVIDIA DGX Spark: one-command install, ~3x upstream prefill, ~1.5x decode, DSpark, and full c...

394
script
script adysec

VPS测试脚本 | VPS性能测试(VPS基本信息、IO性能、全球测速、ping、回程路由测试)、BBR加速脚本(一种加速TCP的拥堵算法技术)、三网测速脚本(三网测速、流媒...

391
InfiniteBench
InfiniteBench OpenBMB Python

Codes for the paper "∞Bench: Extending Long Context Evaluation Beyond 100K Tokens": https://arxiv.org/abs/2402.13718

391
superbenchmark
superbenchmark microsoft Python

A validation and profiling tool for AI infrastructure

391
ETA-Resource
ETA-Resource linwhitehat

Materials about Encrypted Traffic Analysis

391
cob
cob knqyf263 Go

Continuous Benchmark for Go Project

390
recsys-dataset
recsys-dataset otto-de Python

🛍 A real-world e-commerce dataset for session-based recommender systems research.

389
ollama-benchmark
ollama-benchmark aidatatools Python

LLM Benchmark for Throughput via Ollama (Local LLMs)

388