Natural language processing (NLP) is a field of computer science that studies how computers and humans interact. In the 1950s, Alan Turing published an article that proposed a measure of intelligence, now called the Turing test. More modern techniques, such as deep learning, have produced results in the fields of language modeling, parsing, and natural-language tasks.
PyTorch implementation of beam search decoding for seq2seq models
[ACL'20] HAT: Hardware-Aware Transformers for Efficient Natural Language Processing
Fast and customizable text tokenization library with BPE and SentencePiece support
Simple downloader for pre-trained word vectors
chatbot_ner: Named Entity Recognition for chatbots.
Neural Search
Tool for interactive embeddings visualization
Implementation of Dynamic memory networks by Kumar et al. http://arxiv.org/abs/1506.07285
Cybertron: the home planet of the Transformers in Go
[EMNLP 2020] OpenUE: An Open Toolkit of Universal Extraction from Text
My completed solutions for CS224N 2021 & 2019
[ICLR 2023] Code for the paper "Binding Language Models in Symbolic Languages"
This repository contains the code related to Natural Language Processing using python scripting language. All the codes are related to my book entitle...
PyContinual (An Easy and Extendible Framework for Continual Learning)
Meandering In Networks of Entities to Reach Verisimilar Answers
A CoNLL-U parser that takes a CoNLL-U formatted string and turns it into a nested python dictionary.
This repository introduces MentaLLaMA, the first open-source instruction following large language model for interpretable mental health analysis.
Sentiment analysis library for russian language
⚡️ Reality OS for Creators
Repository for Project Insight: NLP as a Service
RASA chatbot use case boilerplate
Natural language processing pipeline for book-length documents (archival Java version; for current Python version, see: https://github.com/booknlp/boo...
NaturalCC: An Open-Source Toolkit for Code Intelligence
Fast and Portable Character String Processing in R (with the Unicode ICU)
ByteNet for character-level language modelling
Mishkal is an arabic text vocalization software
On-device LLM Inference Powered by X-Bit Quantization
Auto get diffusion nlp papers in Axriv. More papers Information can be found in another repository "Diffusion-LM-Papers".
Based on the Pytorch-Transformers library by HuggingFace. To be used as a starting point for employing Transformer models in text classification tasks...
code samples for the goodreads datasets
Default English stopword lists from many different sources
A list of contrastive Learning papers
Research and Materials on Hardware implementation of Transformer Model
Text2Text Language Modeling Toolkit
[ECCV 2020] ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language
a sklearn wrapper for Google's BERT model
this repository features assignments and projects from the iNeuron full stack data science course, providing valuable resources for learners to enhanc...
Open-sourced course notes for Artificial Intelligence and Data Science related topics, prepared in LaTeX
A personal semantic search engine capable of surfacing relevant bookmarks, journal entries, notes, blogs, contacts, and more, built on an efficient do...
Tracking the progress in non-autoregressive generation (translation, transcription, etc.)
An NLP system for generating reading comprehension questions
This package features data-science related tasks for developing new recognizers for Presidio. It is used for the evaluation of the entire system, as w...
AI ChatBot using Python Tensorflow and Natural Language Processing (NLP) along side TFLearn
Recent Deep Learning papers in NLU and RL
Pre-Trained Models for ToD-BERT
All NLP you Need Here. 目前包含15个NLP demo的pytorch实现(大量代码借鉴于其他开源项目,原先是自己玩的,后来干脆也开源出来)
LDA topic modeling for node.js
RETVec is an efficient, multilingual, and adversarially-robust text vectorizer.
An Extensible Continual Learning Framework Focused on Language Models (LMs)
[ICLR 2024] Mol-Instructions: A Large-Scale Biomolecular Instruction Dataset for Large Language Models