Online reviews hold detailed but unstructured evidence about how customers experience a business, yet managers seldom have the time or tooling to convert this text into prioritised and trustworthy actions. This paper presents an AI-driven service analytics framework that transforms raw customer reviews into aspect-level insights and explains the reasoning behind every sentiment judgement. Reviews are encoded with a compact sentence transformer (all-MiniLM-L6-v2) and the resulting dense vectors are kept in a lightweight SQLite repository. For each business aspect, namely food, service, ambience and pricing, a retrieval stage selects the most semantically relevant reviews through cosine similarity, after which a DistilBERT classifier fine-tuned on SST-2 estimates the positive-class probability of the aspect-bearing sentences. The averaged scores are mapped to positive, mixed or negative labels and linked to a rule base of managerial recommendations that is ordered so the weakest aspect is addressed first. To make every output auditable, SHAP partition explanations attribute each prediction to individual tokens. On a balanced sample of restaurant reviews drawn from the UCI Sentiment Labelled Sentences corpus, the framework scored service and ambience at 0.75, food at 0.66 and pricing at 0.50, identifying pricing as the first improvement priority. The SHAP study of a contrastive review showed that the conjunction “but” and negative descriptors counterbalance a strongly positive opening word, demonstrating how explanations reveal the model’s internal reasoning. The entire pipeline runs locally with open-source models and no paid APIs, which makes it practical for small and medium enterprises.
Introduction
The text presents an explainable NLP-based framework for analyzing customer reviews of restaurants. Its main purpose is to convert large amounts of unstructured customer feedback into aspect-level sentiment, prioritized business recommendations, and understandable explanations.
Traditional customer feedback methods such as surveys and star ratings provide limited information. A star rating gives an overall impression but does not explain whether the problem concerns food, service, ambience, or pricing. Similarly, ordinary sentiment analysis can determine whether a review is positive or negative but cannot identify the specific aspect responsible for that sentiment.
To solve this problem, the proposed system combines semantic retrieval, transformer-based sentiment analysis, rule-based decision support, and SHAP explainability.
Proposed Framework
The system consists of several main stages:
Review collection and preprocessing – Customer reviews are collected and prepared for analysis.
Semantic embedding – Reviews are converted into 384-dimensional embeddings using all-MiniLM-L6-v2.
Aspect-oriented retrieval – Reviews relevant to four aspects—food, service, ambience, and pricing—are retrieved using cosine similarity rather than relying only on exact keywords.
Sentiment analysis – Retrieved evidence is analyzed using DistilBERT, which estimates the probability of positive sentiment.
Aspect classification – Scores are converted into three categories:
Positive: score ≥ 0.65
Mixed: 0.40–0.65
Negative: score ≤ 0.40
Decision support – The system converts sentiment results into specific managerial recommendations and ranks aspects according to their need for attention.
Explainability – SHAP identifies the words or tokens that contribute to the model's sentiment prediction.
Dataset and Implementation
The experiment uses the Yelp section of the UCI Sentiment Labelled Sentences dataset. From 1,000 restaurant review sentences, a balanced sample of 50 reviews—25 positive and 25 negative— was selected.
The system was implemented in Python on Google Colab using CPU resources. It uses open-source technologies including Sentence Transformers, DistilBERT, SHAP, NumPy, SQLite, pandas, and Matplotlib. SQLite was selected as a lightweight local database, making the system suitable for small businesses without expensive cloud infrastructure or commercial AI APIs.
Main Results
The framework evaluates four restaurant aspects:
Aspect
Score
Label
Main Interpretation
Pricing
0.50
Mixed
Main area requiring improvement
Food
0.66
Positive
Strength, but relatively close to the threshold
Service
0.75
Positive
Strong aspect
Ambience
0.75
Positive
Strong aspect
The results indicate that pricing was the main area requiring attention, while service and ambience were identified as stronger aspects. Food was classified as positive but was relatively close to the positive threshold.
The retrieval experiment also showed that semantic embeddings could identify food-related reviews even when they did not contain exactly the same keywords as the aspect description.
Explainability
SHAP was used to demonstrate why the sentiment classifier produced its prediction. In the example review:
“Amazing food but terrible slow service and overpriced menu”
words such as “Amazing” contributed positively, while “terrible,” “slow,” and “but” contributed negatively. This demonstrates that the system can expose the linguistic evidence behind its predictions instead of simply producing an unexplained sentiment score.
Research Gap and Contribution
The paper argues that previous research often addresses sentiment analysis, aspect detection, retrieval, or explainability separately. The proposed framework combines these components into one lightweight decision-support system.
Its main contributions are:
Aspect-oriented retrieval without requiring manually labelled aspect terms.
Actionable recommendations based on transparent scoring rules.
SHAP-based explanations showing which words influence predictions.
Low-cost local deployment using open-source models and SQLite.
Limitations
The study has several limitations. The evaluation uses only 50 short reviews from one dataset, so the results cannot establish broad statistical generalization. There is also no aspect-level ground truth, making quantitative evaluation of aspect-level precision and recall difficult. The aspect results depend on the selected descriptors and retrieval value k, while the sentiment model was trained on movie reviews (SST-2) rather than restaurant-specific data. In addition, SHAP took approximately 50 seconds for one sentence on CPU, which could be problematic for real-time applications.
Conclusion
This paper presented an explainable service analytics framework that converts customer reviews into aspect-level sentiment, prioritised recommendations and token-level explanations. By pairing dense semantic retrieval with a distilled transformer classifier, the framework isolates evidence for food, service, ambience and pricing without aspect annotations. A transparent thresholding and rule layer turns scores into actions, and SHAP attributions allow each judgement to be inspected. In the case study, the framework identified pricing as the principal area for improvement while confirming service and ambience as strengths, and its explanations revealed how contrastive words and sub-word fragments shape the model’s output. Because it relies only on open-source components and an embedded database, the solution can be adopted by small and medium enterprises at minimal cost.
Future work will extend the evaluation to larger, multi-business review collections with aspect-level annotations, such as the SemEval restaurant benchmarks, to measure accuracy quantitatively. Domain-adapted sentiment models, automatic discovery of new aspects through clustering, temporal tracking of sentiment trends and an interactive dashboard for managers are further planned enhancements. Integrating a generative language model to draft context-specific recommendations from the retrieved evidence, while retaining SHAP-based verification, is another promising direction.
References
[1] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proc. NAACL-HLT, 2019, pp. 4171–4186.
[2] V. Sanh, L. Debut, J. Chaumond, and T. Wolf, “DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter,” arXiv preprint arXiv:1910.01108, 2019.
[3] M. Pontiki, D. Galanis, J. Pavlopoulos, H. Papageorgiou, I. Androutsopoulos, and S. Manandhar, “SemEval-2014 Task 4: Aspect based sentiment analysis,” in Proc. 8th Int. Workshop on Semantic Evaluation (SemEval 2014), 2014, pp. 27–35.
[4] M. Hu and B. Liu, “Mining and summarizing customer reviews,” in Proc. 10th ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining, 2004, pp. 168–177.
[5] A. Adadi and M. Berrada, “Peeking inside the black-box: A survey on explainable artificial intelligence (XAI),” IEEE Access, vol. 6, pp. 52138–52160, 2018.
[6] A. Barredo Arrieta et al., “Explainable artificial intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI,” Information Fusion, vol. 58, pp. 82–115, 2020.
[7] [7] S. M. Lundberg and S.-I. Lee, “A unified approach to interpreting model predictions,” in Advances in Neural Information Processing Systems, vol. 30, 2017, pp. 4765–4774.
[8] P. Lewis et al., “Retrieval-augmented generation for knowledge-intensive NLP tasks,” in Advances in Neural Information Processing Systems, vol. 33, 2020, pp. 9459–9474.
[9] V. Karpukhin et al., “Dense passage retrieval for open-domain question answering,” in Proc. Conf. Empirical Methods in Natural Language Processing, 2020, pp. 6769–6781.
[10] B. Pang and L. Lee, “Opinion mining and sentiment analysis,” Foundations and Trends in Information Retrieval, vol. 2, no. 1–2, pp. 1–135, 2008.
[11] R. Socher, A. Perelygin, J. Wu, J. Chuang, C. D. Manning, A. Ng, and C. Potts, “Recursive deep models for semantic compositionality over a sentiment treebank,” in Proc. Conf. Empirical Methods in Natural Language Processing, 2013, pp. 1631–1642.
[12] L. Zhang, S. Wang, and B. Liu, “Deep learning for sentiment analysis: A survey,” WIREs Data Mining and Knowledge Discovery, vol. 8, no. 4, e1253, 2018.
[13] A. Vaswani et al., “Attention is all you need,” in Advances in Neural Information Processing Systems, vol. 30, 2017, pp. 5998–6008.
[14] T. Wolf et al., “Transformers: State-of-the-art natural language processing,” in Proc. Conf. Empirical Methods in Natural Language Processing: System Demonstrations, 2020, pp. 38–45.
[15] N. Reimers and I. Gurevych, “Sentence-BERT: Sentence embeddings using Siamese BERT-networks,” in Proc. Conf. Empirical Methods in Natural Language Processing (EMNLP-IJCNLP), 2019, pp. 3982–3992.
[16] W. Wang, F. Wei, L. Dong, H. Bao, N. Yang, and M. Zhou, “MiniLM: Deep self-attention distillation for task-agnostic compression of pre-trained transformers,” in Advances in Neural Information Processing Systems, vol. 33, 2020, pp. 5776–5788.
[17] M. T. Ribeiro, S. Singh, and C. Guestrin, “‘Why should I trust you?’: Explaining the predictions of any classifier,” in Proc. 22nd ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining, 2016, pp. 1135–1144.
[18] D. Kotzias, M. Denil, N. de Freitas, and P. Smyth, “From group to individual labels using deep features,” in Proc. 21st ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining, 2015, pp. 597–606.