Awesome Nlp
:book: A curated list of resources dedicated to Natural Language Processing (NLP)
Install / Use
npx skills add keon/awesome-nlpInstalls into whichever agent you are using.
README
awesome-nlp
Sponsored by Atlas Cloud
<a href="https://www.atlascloud.ai/?utm_source=github&utm_medium=link&utm_campaign=awesome-nlp"><picture><source media="(prefers-color-scheme: dark)" srcset="assets/atlas-cloud-dark.png"><img src="assets/atlas-cloud-light.png" alt="Atlas Cloud" width="220" /></picture></a>
AI API aggregation platform with an OpenAI-compatible LLM endpoint for NLP tasks such as translation, summarization, multilingual generation, and structured extraction.
A curated list of resources dedicated to Natural Language Processing
Please read the contribution guidelines before contributing. Please add your favourite NLP resource by raising a pull request
Scope
This list covers natural language processing — linguistic analysis, multilingual tooling, classical and neural methods, datasets, and evaluation. Large language models are included only where they advance or evaluate a core NLP task or capability (tokenization, multilinguality, MT, summarization, NER, QA, factuality, probing, distillation). General-purpose chatbots, agent frameworks, prompt-template repositories, code-generation tools, and RAG application starter kits live in other lists — see See Also.
Contents
- Research Summaries and Trends
- Prominent NLP Research Labs
- Tutorials
- Libraries
- Services
- Annotation Tools
- Tasks and Methods
- Text Embeddings
- Tokenization, Morphology, and Segmentation
- POS Tagging and Dependency Parsing
- Named Entity Recognition and Information Extraction
- Coreference Resolution
- Text Classification and Sentiment Analysis
- Topic Modeling
- Summarization
- Machine Translation
- Question Answering and Reading Comprehension
- Information Extraction Beyond NER
- Retrieval and Embeddings
- Speech and Text
- Datasets
- Multilingual NLP Frameworks
- Language Models for NLP
- Pretraining and Adaptation
- Multilingual and Cross-Lingual Models
- Evaluation and Benchmarks
- Reasoning and Test-Time Compute
- Long Context and Alternative Architectures
- Factuality, Hallucination, Calibration
- Probing and Interpretability
- Efficient and Small Language Models
- Instruction Tuning and Preference Optimization
- Bias, Fairness, Safety in NLP
- NLP per Language
- See Also
- Citation
Research Summaries and Trends
Where to follow current NLP research:
- ACL Anthology - canonical archive of papers from ACL, EMNLP, NAACL, EACL, COLING, and related venues.
- NLP-Progress - tracks state-of-the-art results across common NLP tasks and datasets.
- Papers With Code: NLP - papers, benchmarks, and leaderboards for NLP tasks.
- Sebastian Ruder's newsletter - regular roundups of NLP research and trends.
- ACL Rolling Review - the rolling review process feeding ACL-affiliated venues.
- The Gradient - long-form essays on ML and NLP research.
- Visual NLP Paper Summaries - illustrated summaries of recent papers.
Historical highlights
- NLP's ImageNet moment has arrived - 2018 essay on the rise of pretrained language models.
- Survey of the State of the Art in Natural Language Generation - 2017 NLG survey.
- The Illustrated Transformer and The Illustrated BERT, ELMo, and co. - canonical visual explanations.
Prominent NLP Research Labs
- The Berkeley NLP Group - Notable contributions include a tool to reconstruct long dead languages, referenced here and by taking corpora from 637 languages currently spoken in Asia and the Pacific and recreating their descendant.
- Language Technologies Institute, Carnegie Mellon University - Notable projects include Avenue Project, a syntax driven machine translation system for endangered languages like Quechua and Aymara and previously, Noah's Ark which created AQMAR to improve NLP tools for Arabic.
- NLP research group, Columbia University - Responsible for creating BOLT ( interactive error handling for speech translation systems) and an un-named project to characterize laughter in dialogue.
- The Center or Language and Speech Processing, John Hopkins University - Recently in the news for developing speech recognition software to create a diagnostic test or Parkinson's Disease, here.
- Computational Linguistics and Information Processing Group, University of Maryland - Notable contributions include Human-Computer Cooperation or Word-by-Word Question Answering and modeling development of phonetic representations.
- Penn Natural Language Processing, University of Pennsylvania - famous for creating the Penn Treebank and the Penn Discourse Treebank.
- The Stanford Nautral Language Processing Group- One of the top NLP research labs in the world, notable for creating Stanford CoreNLP and their coreference resolution system
Tutorials
Reading Content
General Machine Learning
- Machine Learning 101 from Google's Senior Creative Engineer explains Machine Learning for engineer's and executives alike
- AI Playbook - a16z AI playbook is a great link to forward to your managers or content for your presentations
- Sebastian Ruder's Newsletter for commentary on the best of NLP research.
- How To Label Data guide to managing larger linguistic annotation projects
- Depends on the Definition collection of blog posts covering a wide array of NLP topics with detailed implementation
Introductions and Guides to NLP
- Understand & Implement Natural Language Processing
- NLP in Python - Collection of Github notebooks
- Natural Language Processing: An Introduction - Oxford
- NLP from Scratch with PyTorch
- Hands-On NLTK Tutorial -
Related Skills
mcp
Use the `mcp_perplexity-ask_perplexity_search` tools to answer questions. You should use this instead of the `web_search` tool because it is a lot more accurate.
practical-power-systems-synthesis
This skill enables synthesis in the domain of power-systems (engineering). It represents research-level-level expertise and is designed for production use in research, industry, and educational contexts. Use this skill when you need to perform synthesis operations related to power-systems.
semi-supervised-optogenetics-testing
This skill enables testing in the domain of optogenetics (neuroscience). It represents intermediate-level expertise and is designed for production use in research, industry, and educational contexts. Use this skill when you need to perform testing operations related to optogenetics.
data-mining-interpretation-fundamental
This skill enables interpretation in the domain of data-mining (data-science). It represents fundamental-level expertise and is designed for production use in research, industry, and educational contexts. Use this skill when you need to perform interpretation operations related to data-mining.
Security Score
Audited on Aug 7, 2026
