Natural language processing (NLP) is a field of computer science that studies how computers and humans interact. In the 1950s, Alan Turing published an article that proposed a measure of intelligence, now called the Turing test. More modern techniques, such as deep learning, have produced results in the fields of language modeling, parsing, and natural-language tasks.
Benchmark for Thai sentence representation
List of AI Internships
[ICLR 2023] Multimodal Analogical Reasoning over Knowledge Graphs
Extensive tutorials for the Advanced NLP Workshop in Open Data Science Conference Europe 2020. We will leverage machine learning, deep learning and de...
How to encode sentences in a high-dimensional vector space, a.k.a., sentence embedding.
《自然语言处理综论》第三版翻译。
EMNLP'23 survey: a curation of awesome papers and resources on refreshing large language models (LLMs) without expensive retraining.
Multilingual Large Language Models Evaluation Benchmark
Powerful web application that combines Streamlit, LangChain, and Pinecone to simplify document analysis. Powered by OpenAI's GPT-3, RAG enables dynami...
<파이토치로 배우는 자연어 처리>(한빛미디어, 2021)의 소스 코드를 위한 저장소입니다.
Active Learning Awesome Paper
The idea is to calculate the similarity between the resume and the job description and then return the resumes with the highest similarity score.
I created an application which takes in live speech or audio recording as input, converts it into text and displays the relevant Indian Sign Language...
Implementation of GPT from scratch. Design to be lightweight and easy to modify.
Cornell Semantic Parsing Framework
Some recipes of natural language pre-processing
Detect common phrases in large amounts of text using a data-driven approach. Size of discovered phrases can be arbitrary. Can be used in languages oth...
Pytorch implementation of "Graph Convolutional Networks for Text Classification"
📝 A list of pre-trained BERT models for Japanese with word/subword tokenization + vocabulary construction algorithm information
XLNet Extension in TensorFlow
Financial Domain Question Answering with pre-trained BERT Language Model
[EMNLP 2021] LingFeat - A Comprehensive Linguistic Features Extraction ToolKit for Readability Assessment
A framework for human-readable prompt-based method with large language models. Specially designed for researchers. (Deprecated, check out LangChain fo...
[ACL 2023 Findings] FACTUAL dataset, the textual scene graph parser trained on FACTUAL.
A curated list of Artificial Intelligence (AI) Research, tracks the cutting edge trending of AI research, including recommender systems, computer visi...
Discovering New Intents with Deep Aligned Clustering (AAAI 2021)
Conversational Toolkit. An Open-Source Toolkit for Fast Development and Fair Evaluation of Text Generation
Pytorch Implementation of our ACL 2020 Paper "Reasoning with Latent Structure Refinement for Document-Level Relation Extraction"
🤖 聊天机器人示例,定制聊天机器人,聊天机器人语料导入导出
KLUE 데이터를 활용한 HuggingFace Transformers 튜토리얼
Stanford CS224n: Natural Language Processing with Deep Learning, Winter 2020
Code for the EMNLP 2021 Paper "Active Learning by Acquiring Contrastive Examples" & the ACL 2022 Paper "On the Importance of Effectively Adapting Pret...
Chat to your database with AI. An experimental app to test the abilities of LLMs to query SQL databases using natural language.
My personal toolkit for PyTorch development.
NLPCC 2016 微博分词评测项目
Research code for ACL 2020 paper: "Distilling Knowledge Learned in BERT for Text Generation".
This repository contains PyTorch implementation for the baseline models from the paper Utterance-level Dialogue Understanding: An Empirical Study
This repository provides everything to get started with Python for Text Mining / Natural Language Processing (NLP)
:speech_balloon: An English-Persian Dictionary of Computer Science and Artificial Intelligence
CrossNorm and SelfNorm for Generalization under Distribution Shifts, ICCV 2021
A wide variety of research projects developed by the SpokenNLP team of Speech Lab, Alibaba Group.
PyTrial: A Comprehensive Platform for Artificial Intelligence for Drug Development
An unobtrusive Obsidian plugin that quietly processes equations and patterns in real time
[ICLR/AAAI/KDD2026] Open-Source LLM-Based Data Analysis Agents
My personal notes and surveys on DL, CV and NLP papers.
A Sentence Cloze Dataset for Chinese Machine Reading Comprehension (CMRC 2019)
使用 RASA NLU 来构建中文自然语言理解系统(NLU)| Use RASA NLU to build a Chinese Natural Language Understanding System (NLU)
Open source tools for Estonian natural language processing
Библиотека для извлечения статистик из текстов на русском языке.
Accelerated NLP pipelines for fast inference on CPU and GPU. Built with Transformers, Optimum and ONNX Runtime.