Natural language processing (NLP) is a field of computer science that studies how computers and humans interact. In the 1950s, Alan Turing published an article that proposed a measure of intelligence, now called the Turing test. More modern techniques, such as deep learning, have produced results in the fields of language modeling, parsing, and natural-language tasks.
Neural Search
AI-powered legal compliance assistant for alcohol beverage pricing laws — extracts, analyzes, and explains New York state-level regulationsusing RAG +...
This repository contains a new generative model of chatbot based on seq2seq modeling.
Cybertron: the home planet of the Transformers in Go
Buddhist Digital Text Platform — 10,500+ texts, 613 sources, trilingual cross-canon, AI Q&A (RAG), knowledge graph, full-text search
Language sentiment analysis and neural networks... for trolls.
[EMNLP 2020] OpenUE: An Open Toolkit of Universal Extraction from Text
dialogbot, provide search-based dialogue, task-based dialogue and generative dialogue model. 对话机器人,基于问答型对话、任务型对话、聊天型对话等模型...
My Tech Blog: about Mojo / Rust / Golang / Python / Kotlin / Flutter / VueJS / Blockchain etc.
A comprehensive Rust translation of the code from Sebastian Raschka's Build an LLM from Scratch book.
A simple effective ToolKit for short text matching
Make function calling with LLM easier
Various models and code (Manhattan LSTM, Siamese LSTM + Matching Layer, BiMPM) for the paraphrase identification task, specifically with the Quora Que...
Interface to ChatGPT from R
This repo is all the machine learning related project codes and their corresponding blog posts at the graduate level.
Machine Translation for Africa
任何 JS 环境可用的中文分词包,fork from leizongmin/node-segment
✨ Bootstrap annotation with zero- & few-shot learning via OpenAI GPT-3
✨ A synthetic dataset generation framework that produces diverse coding questions and verifiable solutions - all in one framwork
Repository for Project Insight: NLP as a Service
AI Tool for querying natural language on tabular data.
Use LLMs to robustly extract web data
自然语言处理项目,目标是对文本进行分类。
Korean Morphological Analyzer by shineware
Multimodal Question Answering in the Medical Domain: A summary of Existing Datasets and Systems
Tutorial for International Summer School on Deep Learning, 2019
Fast and Portable Character String Processing in R (with the Unicode ICU)
PYthon Automated Term Extraction
BERT distillation(基于BERT的蒸馏实验 )
Edge Inference in Browser with Transformer NLP model
⛵️The official PyTorch implementation for "BERT-of-Theseus: Compressing BERT by Progressive Module Replacing" (EMNLP 2020).
中文智能客服机器人demo,包含闲聊和专业问答2个部分,支持自定义组件(Chinese intelligent customer chatbot Demo, including the gossip and the professiona...
Package to compute Mauve, a similarity score between neural text and human text. Install with `pip install mauve-text`.
A powerful dataset generator for Rasa NLU, inspired by Chatito
2019百度的关系抽取比赛,使用Pytorch实现苏神的模型,F1在dev集可达到0.75,联合关系抽取,Joint Relation Extraction.
Persian Swear Dataset - you can use in your production to filter unwanted content. دیتاست کلمات نامناسب و بد فارسی برای فیلتر کردن متن ها
Links to Russian corpora + Python functions for loading and parsing
AI-writing humanizer for KO/EN/ZH/JA
Default English stopword lists from many different sources
啊哈自然语言处理包,提供包括分词、依存句法分析、语义角色标注、自动摘要、语义相似度计算、LDA 主题预测、词云等服务。
Codes for ICML 2022 paper: Matching Structure for Dual Learning
VeritasGraph — open-source Knowledge Graph & GraphRAG framework on GitHub. Build multi-hop reasoning, ontology-aware retrieval, and verifiable attribu...
multilabel classification of EHR notes
BNLP is a natural language processing toolkit for Bengali Language.
AubAI brings you on-device gen-AI capabilities, including offline text generation and more, directly within your app.
A curated list of resources dedicated to Natural Language Processing (NLP) in polish. Models, tools, datasets.
A joint community effort to create one central leaderboard for LLMs.
JAX library for training sub-4B foundation models for edge
A pytorch implementation of the ACL2019 paper "Simple and Effective Text Matching with Richer Alignment Features".
Minecraft-style voxel benchmark for comparing AI models (Arena + Sandbox)