Well-established evaluation metrics, such as F1 score and syntactic graph properties, are routinely used in computational linguistics to measure the effectiveness and accuracy of parsing algorithms.
The retrieved literature consistently demonstrates the use of established evaluation metrics, accuracy scores, and complexity analyses to measure the effectiveness of various parsing algorithms (e.g., PCFGs, dependency parsers, and tree-averaging algorithms). None of the papers refute this claim.