Natural language processing (NLP) is a field of computer science that studies how computers and humans interact. In the 1950s, Alan Turing published an article that proposed a measure of intelligence, now called the Turing test. More modern techniques, such as deep learning, have produced results in the fields of language modeling, parsing, and natural-language tasks.
Code for the paper "Are Sixteen Heads Really Better than One?"
Unilm for Chinese Chitchat Robot.基于Unilm模型的夸夸式闲聊机器人项目。
Content Enhanced BERT-based Text-to-SQL Generation https://arxiv.org/abs/1910.07179
Solution to Kaggle's Quora Duplicate Question Detection Competition
MATILDA: Multi-AnnoTator multi-language Interactive Lightweight Dialogue Annotator
A baseline implementation for FNC-1
Convert number words (eg. twenty one) to numeric digits (21)
Natural language detection package in pure Go
Stanford NLP group's shared Python tools.
A PyTorch implementation of Mnemonic Reader for the Machine Comprehension task
Python package containing all custom layers used in Neural Networks (Compatible with PyTorch, TensorFlow and MegEngine)
A fast and accurate POS and morphological tagging toolkit (EACL 2014)
Gitbook Address: https://app.gitbook.com/@nlpgroup/s/nlpnote/
Lightweight, Python library for fast and reproducible experimentation :microscope:
Quickly turn command-line applications into RESTful webservices with a web-application front-end. You provide a specification of your command line app...
A curated list of Clojure resources for dealing with domain-specific languages.
Natural Language Processing For Everyone
Python wrapper for Stanford CoreNLP's SUTime
Source codes and corpora of paper "Iterated Dilated Convolutions for Chinese Word Segmentation"
An example for applying FusionNet to Natural Language Inference
🇨🇳🇬🇧Chinese and English word spelling corrector.(中文易错别字检测,中文拼写检测纠正。英文单词拼写校验工具)
Educational material on using the TensorFlow Estimator framework for text classification
瑞金医院MMC人工智能辅助构建知识图谱大赛初赛
Detect common phrases in large amounts of text using a data-driven approach. Size of discovered phrases can be arbitrary. Can be used in languages oth...
📝 A list of pre-trained BERT models for Japanese with word/subword tokenization + vocabulary construction algorithm information
TensorFlow implementation of Match-LSTM and Answer pointer for the popular SQuAD dataset.
bert chinese similarity
:newspaper: High-performance tool for negation and uncertainty detection in radiology reports
The official implementation of ACL 2019 paper "Topic-Aware Neural Keyphrase Generation for Social Media Language"
NLPGym - A toolkit to develop RL agents to solve NLP tasks.
List of textual data sources to be used for text mining in R
:smile: Dataset for Emotion Classification
Pytorch implementation of Paragraph-level Neural Question Generation with Maxout Pointer and Gated Self-attention Networks
aim to use JapaneseTokenizer as easy as possible
Load GPT-2 checkpoint and generate texts
CRFSharp is Conditional Random Fields implemented by .NET(C#), a machine learning algorithm for learning from labeled sequences of examples.
Implementation of Hierarchical Attention Networks in PyTorch
Python wrapper for wit.ai's Duckling Clojure library
Simple NLP in Rust with Python bindings
Turkish deasciifier in Python based on Deniz Yüret's turkish-mode for Emacs
💫 Scripts, tools and resources for developing spaCy
Morphological analyzer for Russian and English languages based on neural networks and dictionary-lookup systems.
Workshop (6 hours): Deep learning in R using Keras. Building & training deep nets, image classification, transfer learning, text analysis, visualizati...
J.A.R.V.I.S - Just Another Rudimentary Verbal Instruction Shell
띄어쓰기 오류 교정 라이브러리입니다. CRF 와 같은 머신러닝 알고리즘이 아닌, 직관적인 접근법으로 띄어쓰기를 교정합니다.
Framework to learn Named Entity Recognition models without labelled data using weak supervision.
Code for "Controllable Unsupervised Text Attribute Transfer via Editing Entangled Latent Representation" (NeurIPS 2019)
Text classification with Convolution Neural Networks on Yelp, IMDB & sentence polarity dataset v1.0
A tool for comparing tokenizers
NLP papers applicable to financial markets