飞桨大模型开发套件,提供大语言模型、跨模态大模型、生物计算大模型等领域的全流程开发工具链。
Swift benchmark runner with many performance metrics and great CI support
Official repository for the ProteinGym benchmarks
OpenML AutoML Benchmarking Framework
A GPU benchmark tool for evaluating GPUs and CPUs on mixed operational intensity kernels (CUDA, OpenCL, HIP, SYCL, OpenMP)
A list of works on evaluation of visual generation models, including evaluation metrics, models, and systems
OWASP VulnerableApp Project: Break it. Scan it. Reproduce it. Benchmark against it. Improve it.
This is a reposotory that includes paper、code and datasets about domain generalization-based fault diagnosis and prognosis. (基于领域泛化的故障诊断和...
Image classification with NVIDIA TensorRT from TensorFlow models.
MCPMark is a comprehensive, stress-testing MCP benchmark designed to evaluate model and agent capabilities in real-world MCP use.
Benchmarking Legal Knowledge of Large Language Models
Benchmark the performances of various Swift layout frameworks (autolayout, UIStackView, PinLayout, LayoutKit, FlexLayout, Yoga, ...)
(Crafter + NetHack) in JAX. ICML 2024 Spotlight.
A universal flight control tuning framework
A cautious, explainable AI-like text risk analyzer for local workflows and coding agents.
🔥 Stupid Simple CPU/MEM "Profiler" for your JS code.
PHP Profiler & Developer Toolbar (built for Phalcon)
SB(SRS Bench) is a set of benchmark and regression test tools, for SRS and other media servers, supports HTTP-FLV, RTMP, HLS, WebRTC and GB28181.
The First Dynamic Map Removal Benchmark | Included 8 SOTA methods | Continous updating
Benchmarks of JavaScript Package Managers
Awesome Deep Learning for Time-Series Imputation, including an unmissable paper and tool list about applying neural networks to impute incomplete time...
This is a simple App to test some blur algorithms on their visual quality and performance.
Chinese Biomedical Language Understanding Evaluation benchmark (ChineseBLUE)
Inspect and Debug your Tauri applications in style 💃
Gym Electric Motor (GEM): An OpenAI Gym Environment for Electric Motors
RoNIN: Robust Neural Inertial Navigation in the Wild
FedScale is a scalable and extensible open-source federated learning (FL) platform.
Benchmarks for the Optimal Power Flow Problem
🔍 The hardest search benchmark in the wild — vague, multi-turn, proactive. 200 long-horizon tasks with persona-driven progressive disclosure, scored...
Database Benchmarking Framework
Are We Fast Yet? Comparing Language Implementations with Objects, Closures, and Arrays
Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live...
PyAF is an Open Source Python library for Automatic Time Series Forecasting built on top of popular pydata modules.
[ICLR'25] OpenRCA: Can Large Language Models Locate the Root Cause of Software Failures?
An extensive evaluation and comparison of 28 state-of-the-art superpixel algorithms on 5 datasets.
CRUD-RAG: A Comprehensive Chinese Benchmark for Retrieval-Augmented Generation of Large Language Models
Remove unwanted files and directories from your node_modules folder
Generate benchmarks for terminal emulators
Framework to evaluate peformance of ROS 2
SKAB - Skoltech Anomaly Benchmark. Time-series data for evaluating Anomaly Detection algorithms.
Jetson Benchmark
GAP Benchmark Suite
Entrpi/ds4, a Blackwell CUDA perf fork of antirez/ds4 on NVIDIA DGX Spark: one-command install, ~3x upstream prefill, ~1.5x decode, DSpark, and full c...
VPS测试脚本 | VPS性能测试(VPS基本信息、IO性能、全球测速、ping、回程路由测试)、BBR加速脚本(一种加速TCP的拥堵算法技术)、三网测速脚本(三网测速、流媒...
Codes for the paper "∞Bench: Extending Long Context Evaluation Beyond 100K Tokens": https://arxiv.org/abs/2402.13718
A validation and profiling tool for AI infrastructure
Materials about Encrypted Traffic Analysis
Continuous Benchmark for Go Project
🛍 A real-world e-commerce dataset for session-based recommender systems research.
LLM Benchmark for Throughput via Ollama (Local LLMs)