I audited all 146 of them. A third were gaps in the reference set and a third were real relationships anchored to things that should not have been entities. Neither is fixed by touching the model.
I write here mostly about evaluation: how measurements on these systems go wrong, and what stops it happening twice.
Writing
Two thirds of my false positives were not extraction errors
A wrong number is usually a scope error, not an arithmetic one
I reported a recall of 0.699. It was 0.641. Nothing had been miscalculated, and by the time I found it the figure had already reached the slides.
Elsewhere