Explosion builds developer tools for AI, Machine Learning and Natural Language Processing.

Project

Topics

Category

Tasks

Select...Code Generation, Coreference Resolution, Dependency Parsing, Distillation, Embeddings & Vectors, Entity Linking, Evaluation, Image Classification, Image Segmentation, Layout Analysis, Lemmatization, Named Entity Recognition, Object Detection, Optical Character Recognition (OCR), Part-of-Speech Tagging, PII Anonymization, Question Answering, Relation Extraction, Retrieval-Augmented Generation (RAG), Rule-Based Matching, Span Categorization, Text Classification, Text Generation, Tokenization, Weak Supervision

Authors

Adriane Boyd, Ákos Kádár, Basile Dura, Chung-Fan Tsai, Damian Romero, Daniel de Kok, Duygu Altinok, Edward Schmuhl, Helena Steckmeister, India Kerle, Ines Montani, Kabir Khan, Lj Miranda, Madeesh Kannan, Magdalena Anioł, Matthew Honnibal, Paul O’Leary McCann, Peter Baumgartner, Philip Vollet, Raphael Mitsch, Rehan Ahmed, Richard Hudson, Ryan Wesslen, Sofie Van Landeghem, Victoria Slocum, Vincent D. Warmerdam, Vinit Ravishankar, Walter Henry

Beta test our upcoming product for agentic NLP

We’re looking for beta partners for Ellf: a platform and virtual assistant that makes your coding agent like Claude Code proficient at developing NLP solutions. Contact us or sign up for the waitlist if your team needs help with a project focused on tasks like information extraction!

RiCoRecA: rich cooking recipe annotation schema

Ventirozos, Jacob-Romero, Alrdahi, Clinch, Batista-Navarro (2026)
The annotation process consists of two sections. Firstly, the annotator utilized a customized Prodigy interface to complete the NER and RC annotation tasks.

Conquering PDFs: document understanding beyond plain text

In this talk, Ines presents a new and modular approach for building robust document understanding systems, using state-of-the-art models and the awesome Python ecosystem.

Feminist AI LAN Party

Three days of workshops, hacking, creating, publishing and connecting locally, featuring a data development workshop with Prodigy and a session on hacking LLMs.

How to advocate for modular NLP in the age of Generative AI

With all the hype around Generative AI, many are led to believe it’s the solution to everything. So how can you, as a developer, communicate the nuances and advocate for new and modular solutions that are better, easier and cheaper?

What the history of the web can teach us about the future of AI

How will AI development look in the future? There is a lot we can learn from another groundbreaking technology: the web. This blog post takes a look at what the history of the web can teach us, and what this means for developers, models, open source and regulation.

How GitLab uses spaCy to analyze support tickets and empower their community

A case study on GitLab’s large-scale NLP pipelines for extracting actionable insights from support tickets and usage questions.

spaCy Chunks v0.0.2

spaCy extension and pipeline component for generating overlapping chunks of sentences or tokens from a document.

Engineering a human-aligned LLM evaluation workflow with Prodigy and DSPy

This post demonstrates a human-in-the-loop workflow for developing and evaluating LLMs, using Prodigy and DSPy to create task-specific, human-aligned metrics that guide model optimization beyond generic evaluation measures.

Accelerate your Career with Open-Source AI

Panel discussion about making a career out of open-source software, featuring Gael Varoquaux (scikit-learn), Steeve Morin (ZML) and Ines.

[Note: Numerous additional blog posts, talks, and research outputs are provided above. Use links to navigate to specific content as needed.]