Natural Language Processing for the next decade. Tokenization, Part-of-Speech Tagging, Named Entity Recognition, Syntactic & Semantic Dependency Parsing, Document Classification
中文分词
An extremely fast implementation of Aho Corasick algorithm based on Double Array Trie.
CS224n: Natural Language Processing with Deep Learning Assignments Winter, 2017
Simple Solution for Multi-Criteria Chinese Word Segmentation
HanLP中文分词Lucene插件,支持包括Solr在内的基于Lucene的系统
Python scripts preprocessing Penn Treebank and Chinese Treebank
Source codes and corpora of paper "Iterated Dilated Convolutions for Chinese Word Segmentation"