May 10, 2026
Evaluation and Error Analysis on a Multitask Review Triage Capstone
The evaluation and error analysis workstream for a GWU graduate capstone project focused on a BERT-based triage classifier with five classification heads.
My graduate capstone project at The George Washington University was a team effort focused on developing a model that analyzes customer reviews and prioritizes those requiring human attention. The model was trained on 12,000 Uber reviews and uses a BERT encoder with five auxiliary heads that feed into a four class triage classifier. This allows support teams to identify urgent safety complaints that might otherwise be overlooked among thousands of routine reviews. My teammates were responsible for the model architecture and labeling strategy. My role focused on evaluation and error analysis. I measured the model’s performance, examined its confusion matrix, and identified the conditions under which it failed. Given the eight-week development timeline, we recognized that the model would have limitations. Rather than simply demonstrating that we had built a classifier, we wanted to establish how well it performed, where it succeeded, and where it failed.

The code and complete analysis notebook are available on GitHub.
Model Architecture
The shared encoder processes each review and passes its ‘[CLS]’ embedding to five task-specific heads. Four heads predict different aspects of the review, including sentiment, emotion, topic, and intent. The fifth head produces an urgency score. These predictions are then combined with the original embedding and passed to a final classification head. The final head assigns each review to one of four priority tiers, ranging from no triage needed to immediate attention.
Results
I evaluated the model across every variant the team tested, including different loss strategies, encoder choices, feature fusion configurations, and auxiliary head ablations. For each variant, I calculated precision, recall, and F1 for every class. I also reported macro and weighted F1, confusion matrices, and urgency error. The final model achieved 82 percent accuracy and a weighted F1 of 0.83 on a held-out set of 1,200 reviews. The difference between weighted and macro F1 reflects the class imbalance in the data. The no triage tier performed nearly perfectly, while most errors occurred in the more closely grouped middle tiers.
The confusion matrix above is generated directly from the notebook data rather than presented as a static image. The notebook reconstructs the matrix, recomputes accuracy and F1, and verifies the results against the reported values, reproducing 0.8225 accuracy, 0.7028 macro F1, and 0.8260 weighted F1. This makes the reported results directly reproducible rather than figures that the reader is simply asked to trust.
Error Analysis
97 percent of the model’s misclassifications occur between adjacent priority tiers. The model rarely confuses a review requiring no triage with one requiring immediate attention, and the most costly type of error is also the rarest. Most confusion occurs between adjacent tiers because the reviews use similar language. Grouping reviews by priority tier and comparing their TF-IDF centroids shows that the urgent and immediate tiers have a cosine similarity of 0.87, while the no triage tier is clearly separated from the higher priority tiers. The confusion matrix and the language of the reviews tell the same story about where the classification boundary becomes difficult.

Label Quality Audit
The model was trained on labels generated in two different ways. One came from a heuristic, while the other came from a separate language model pass. I treated the two labeling methods as independent annotators and measured the level of agreement between them. They agreed only moderately, with a Cohen’s kappa of approximately 0.33. The disagreement was also directional. The heuristic escalated reviews far more aggressively than the model did. This noise in the training labels places a real limit on how confidently the triage accuracy can be interpreted. It is also the kind of issue that should be identified before placing significant weight on the model’s results.
Availability
This project is an analysis notebook rather than a deployed application, so there is no live application to launch. The notebook runs on a CPU using the public review data and completes in a few seconds. The rendered version linked above presents each figure and result alongside the code that produced it. The source code and complete notebook are available on GitHub.
Receipts
- Dataset 12,000 Uber app store reviews across four triage tiers.
- Result 0.83 weighted F1 and 0.70 macro F1. 97.2% of errors occur between adjacent tiers.
- Label audit Cohen's κ between the two labeling pipelines is reported in the README. This disagreement places a meaningful limit on how confidently the reported metric can be interpreted.
- Source github.com/rajeshnandipaty/review-triage