💫 Improve annotation serialisation #1045

honnibal · 2017-05-07T16:11:14Z

A persistent source of problems in spaCy 1.0 has been the way data is saved and loaded. This issue describes what's changing, and will be updated as implementation proceeds. The changes will finally allow annotations to support the Pickle protocol, making it much easier to use spaCy with Spark and other tools.

How saving a `Doc` works in 1.x

Annotations are exported as integer IDs, into a numpy array. A custom Huffman-tree implementation is used to store the following fields:

ORTH –- backing off to characters for unseen words
SPACY (boolean flag for whether the word has a space after it)
HEAD (as offset from token)
TAG (part of speech tag, must have entry in tag map)
ENT_IOB (IOB format for entity tags – one of 0, 1, 2, 3)
ENT_TYPE (Must have entry in strings.json)
DEP (Dependency label – must have entry in strings.json)

To load the serialized bytes into a new Doc object, you need a matching Vocab: to decode the bytes into the correct integer IDs, and then to map the integer IDs to the correct lexemes and tags.

How saving a `Doc` will work in 2.x

Individual documents will now be serialized as a tuple (attrs, text). Instead of our own Huffman codec, we'll just store the numpy arrays directly –- these compress fine in bulk anyway, especially when merged.

To save and load the document, you'll still need to have a reference to the Vocab object. To assist this, a new Binder class will be introduced, which will allow a group of documents to be saved and loaded together.

The Binder will support Pickle and JSON serialisation. More serialisation protocols can be added in future.

The Binder will support two header styles: a standalone format, and a diff against a model ID. If you serialise the Binder standalone, you'll read out the whole vocab, so this will be large for a small set of documents. However, it will ensure you'll be able to load the documents correctly in future. The diff format will require you to have the appropriate model ID loaded. The Binder will then include in its header any extra information not in the base model, for instance missing strings.

Summary of code changes

New spacy.tokens.binder module, with spacy.tokens.binder.Binder class
Support for Pickle protocol in Doc Token, Span objects
Remove spacy.serialize subpackage

Related issues

The text was updated successfully, but these errors were encountered:

ines · 2017-06-05T23:45:04Z

See the v2.0.0 alpha release notes and #1105 🎉

lock · 2018-05-08T20:38:55Z

This thread has been automatically locked since there has not been any recent activity after it was closed. Please open a new issue for related bugs.

ines added enhancement Feature requests and improvements 🌙 nightly Discussion and contributions related to nightly builds ⚠️ wip Work in progress labels May 7, 2017

ines closed this as completed Jun 5, 2017

stefah mentioned this issue Aug 17, 2017

[2.0] Using spaCy in combination with Apache Spark; still can't be pickled? #1270

Closed

lock bot locked as resolved and limited conversation to collaborators May 8, 2018

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

💫 Improve annotation serialisation #1045

💫 Improve annotation serialisation #1045

honnibal commented May 7, 2017 •

edited by ines

Loading

ines commented Jun 5, 2017

lock bot commented May 8, 2018

💫 Improve annotation serialisation #1045

💫 Improve annotation serialisation #1045

Comments

honnibal commented May 7, 2017 • edited by ines Loading

How saving a Doc works in 1.x

How saving a Doc will work in 2.x

Summary of code changes

Related issues

ines commented Jun 5, 2017

lock bot commented May 8, 2018

honnibal commented May 7, 2017 •

edited by ines

Loading

How saving a `Doc` works in 1.x

How saving a `Doc` will work in 2.x