Matcher is too slow processing text #1819

sandeep118 · 2018-01-09T19:54:32Z

I built a matcher & phrase matcher object which has 10 lakh patterns. For some reason Matcher is taking very long time to process a document ( > 5 mins for each doc) and phrase matcher is doing it in a second on the same data.

I don't know if I am missing something here, but any help is much appreciated.

The reason I need to use Matcher instead of phrase matcher because:

Ex: I have a text file with many phrases such as "United State Of America"
I built phrase matcher object using the above. In the test document if there is phrase
"United State Of \n America" phrase matcher will not return a match because of the extra newline character. Whereas in Matcher I can customize the pattern by adding various things like {'ORTH':'\n','OP':'*'}.

Your Environment

spaCy version: 2.0.5
Platform: Darwin-17.2.0-x86_64-i386-64bit
Python version: 3.6.0
Models: en_core_web_lg

ines · 2018-02-12T11:12:09Z

Merging this with the master issue #1971!

lock · 2018-05-07T23:55:03Z

This thread has been automatically locked since there has not been any recent activity after it was closed. Please open a new issue for related bugs.

honnibal added the performance label Jan 12, 2018

ines mentioned this issue Feb 12, 2018

💫 Better, faster and more customisable matcher #1971

Closed

5 tasks

ines closed this as completed Feb 12, 2018

lock bot locked as resolved and limited conversation to collaborators May 7, 2018

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Matcher is too slow processing text #1819

Matcher is too slow processing text #1819

sandeep118 commented Jan 9, 2018 •

edited

Loading

ines commented Feb 12, 2018

lock bot commented May 7, 2018

Matcher is too slow processing text #1819

Matcher is too slow processing text #1819

Comments

sandeep118 commented Jan 9, 2018 • edited Loading

Your Environment

ines commented Feb 12, 2018

lock bot commented May 7, 2018

sandeep118 commented Jan 9, 2018 •

edited

Loading