Explosion builds developer tools for AI, Machine Learning and Natural Language Processing.
Project
Topics
Category
Tasks
Select...Code Generation, Coreference Resolution, Dependency Parsing, Distillation, Embeddings & Vectors, Entity Linking, Evaluation, Image Classification, Image Segmentation, Layout Analysis, Lemmatization, Named Entity Recognition, Object Detection, Optical Character Recognition (OCR), Part-of-Speech Tagging, PII Anonymization, Question Answering, Relation Extraction, Retrieval-Augmented Generation (RAG), Rule-Based Matching, Span Categorization, Text Classification, Text Generation, Tokenization, Weak Supervision.
Authors
Select...Adriane Boyd, Ákos Kádár, Basile Dura, Chung-Fan Tsai, Damian Romero, Daniel de Kok, Duygu Altinok, Edward Schmuhl, Helena Steckmeister, India Kerle, Ines Montani, Kabir Khan, Lj Miranda, Madeesh Kannan, Magdalena Anioł, Matthew Honnibal, Paul O’Leary McCann, Peter Baumgartner, Philip Vollet, Raphael Mitsch, Rehan Ahmed, Richard Hudson, Ryan Wesslen, Sofie Van Landeghem, Victoria Slocum, Vincent D. Warmerdam, Vinit Ravishankar, Walter Henry.
To create the infrastructure needed for supporting this effort, we repurpose an existing commercial annotation tool, Prodigy, which we then enhance with additional components for combining the annotation tasks into pipelines, cross-validating the annotations and supporting annotator access to tasks.
📚 spacy-layout v0.0.12 Mar 8, 2025 \nSupport processing PDFs with context, add document index tables and more docs
We use the spaCy library for tokenization, part-of-speech tagging, and lemmatization of the words in the descriptions.
In this video, we’ll show how you might be able to improve the annotation experience by using bulk labelling for image classification.
In this talk, Ines presents a new and modular approach for building robust document understanding systems, using state-of-the-art models and the awesome Python ecosystem.
Context aware, pluggable and customizable PII de-identification and anonymization service for text and images, featuring a spaCy back-end.
In this video, we’ll show you how to use Prodigy to find bad examples in the Google QuickDraw dataset. We will be leveraging a technique that involves UMAP to find strange images semi-automatically.
In this video, we’ll show you how you can use Prodigy to script fully custom annotation workflows in Python, how to plug in your own machine learning models and how to mix and match different interfaces for your specific use case.
This blog post presents a new modular workflow for converting PDFs and similar documents to structured data and shows you how to build end-to-end document understanding and information extraction pipelines for industry use cases.