Explosion builds **developer tools** for AI, Machine Learning and Natural Language Processing.

### Project

- [spaCy](/content/_/project/spacy/index.html)
- [Prodigy](/content/_/project/prodigy/index.html)
- [Ellf](/content/_/project/ellf/index.html)
- [Thinc](/content/_/project/thinc/index.html)

### Topics

- [LLMs](/content/_/topic/llms/index.html)
- [NLP Strategy](/content/_/topic/strategy/index.html)
- [Annotation](/content/_/topic/annotation/index.html)
- [Biomedical](/content/_/topic/biomedical/index.html)
- [Finance](/content/_/topic/finance/index.html)
- [Media](/content/_/topic/media/index.html)
- [Legal](/content/_/topic/legal/index.html)
- [Humanities](/content/_/topic/humanities/index.html)
- [Computer Vision](/content/site-root.html)

### Category

- [Blog](/content/_/category/blog/index.html)
- [Release](/content/_/category/release/index.html)
- [Universe](/content/_/category/universe/index.html)
- [Talk](/content/_/category/talk/index.html)
- [Interview](/content/_/category/interview/index.html)
- [Video](/content/_/category/video/index.html)
- [Paper](/content/_/category/paper/index.html)
- [Book](/content/_/category/book/index.html)

### Tasks

Select...Code Generation, Coreference Resolution, Dependency Parsing, Distillation, Embeddings & Vectors, Entity Linking, Evaluation, Image Classification, Image Segmentation, Layout Analysis, Lemmatization, Named Entity Recognition, Object Detection, Optical Character Recognition (OCR), Part-of-Speech Tagging, PII Anonymization, Question Answering, Relation Extraction, Retrieval-Augmented Generation (RAG), Rule-Based Matching, Span Categorization, Text Classification, Text Generation, Tokenization, Weak Supervision.

### Authors

Select...Adriane Boyd, Ákos Kádár, Basile Dura, Chung-Fan Tsai, Damian Romero, Daniel de Kok, Duygu Altinok, Edward Schmuhl, Helena Steckmeister, India Kerle, Ines Montani, Kabir Khan, Lj Miranda, Madeesh Kannan, Magdalena Anioł, Matthew Honnibal, Paul O’Leary McCann, Peter Baumgartner, Philip Vollet, Raphael Mitsch, Rehan Ahmed, Richard Hudson, Ryan Wesslen, Sofie Van Landeghem, Victoria Slocum, Vincent D. Warmerdam, Vinit Ravishankar, Walter Henry.

## [Building Multimodal Corpora Using Microtask Pipelines and Local Annotators](https://lrec.elra.info/lrec2026-main-514) [Hotti, Vázquez, Jokipohja, Kalliokoski, Paakki, Suviranta, Hiippala (2026)](https://lrec.elra.info/lrec2026-main-514)

[To create the infrastructure needed for supporting this effort, we repurpose an existing commercial annotation tool, Prodigy, which we then enhance with additional components for combining the annotation tasks into pipelines, cross-validating the annotations and supporting annotator access to tasks.](https://lrec.elra.info/lrec2026-main-514)

[**📚 spacy-layout v0.0.12 Mar 8, 2025** \nSupport processing PDFs with context, add document index tables and more docs](https://github.com/explosion/spacy-layout)

## [Describing Images Fast and Slow: Quantifying and Predicting the Variation in Human Signals during Visuo-Linguistic Processes](https://arxiv.org/abs/2402.01352) [Takmaz, Pezzelle, Fernández (2024)](https://arxiv.org/abs/2402.01352)

[We use the spaCy library for tokenization, part-of-speech tagging, and lemmatization of the words in the descriptions.](https://arxiv.org/abs/2402.01352)

## [Finetuning and Bulk Labelling Images with Prodigy](https://www.youtube.com/watch?v=DmH3JmX3w2I)

[In this video, we’ll show how you might be able to improve the annotation experience by using bulk labelling for image classification.](https://www.youtube.com/watch?v=DmH3JmX3w2I)

## [Conquering PDFs: document understanding beyond plain text](https://speakerdeck.com/inesmontani/conquering-pdfs-document-understanding-beyond-plain-text) [PyData London](https://speakerdeck.com/inesmontani/conquering-pdfs-document-understanding-beyond-plain-text)

[In this talk, Ines presents a new and modular approach for building robust document understanding systems, using state-of-the-art models and the awesome Python ecosystem.](https://speakerdeck.com/inesmontani/conquering-pdfs-document-understanding-beyond-plain-text)

## [Microsoft Presidio v2.2.352](https://github.com/microsoft/presidio)

[Context aware, pluggable and customizable PII de-identification and anonymization service for text and images, featuring a spaCy back-end.](https://github.com/microsoft/presidio)

## [Finding Bad Image Data using UMAP and Prodigy](https://www.youtube.com/watch?v=s0Y45xscE-0)

[In this video, we’ll show you how to use Prodigy to find bad examples in the Google QuickDraw dataset. We will be leveraging a technique that involves UMAP to find strange images semi-automatically.](https://www.youtube.com/watch?v=s0Y45xscE-0)

## [Image Captioning with Prodigy & PyTorch](https://www.youtube.com/watch?v=zlyq9z7hdUA)

[In this video, we’ll show you how you can use Prodigy to script fully custom annotation workflows in Python, how to plug in your own machine learning models and how to mix and match different interfaces for your specific use case.](https://www.youtube.com/watch?v=zlyq9z7hdUA)

## [From PDFs to AI-ready structured data: a deep dive](/content/blog/pdfs-nlp-structured-data/index.html)

[This blog post presents a new modular workflow for converting PDFs and similar documents to structured data and shows you how to build end-to-end document understanding and information extraction pipelines for industry use cases.](/content/blog/pdfs-nlp-structured-data/index.html)
