Fake consumer review detection using deep neural networks integrating word embeddings and emotion mining

Hajek, Petr; Barushka, Aliaksandr; Munk, Michal

doi:10.1007/s00521-020-04757-2

Fake consumer review detection using deep neural networks integrating word embeddings and emotion mining

S.I. : Emerging applications of Deep Learning and Spiking ANN
Published: 01 February 2020

Volume 32, pages 17259–17274, (2020)
Cite this article

Neural Computing and Applications Aims and scope Submit manuscript

4687 Accesses
88 Citations
Explore all metrics

Abstract

Fake consumer review detection has attracted much interest in recent years owing to the increasing number of Internet purchases. Existing approaches to detect fake consumer reviews use the review content, product and reviewer information and other features to detect fake reviews. However, as shown in recent studies, the semantic meaning of reviews might be particularly important for text classification. In addition, the emotions hidden in the reviews may represent another potential indicator of fake content. To improve the performance of fake review detection, here we propose two neural network models that integrate traditional bag-of-words as well as the word context and consumer emotions. Specifically, the models learn document-level representation by using three sets of features: (1) n-grams, (2) word embeddings and (3) various lexicon-based emotion indicators. Such a high-dimensional feature representation is used to classify fake reviews into four domains. To demonstrate the effectiveness of the presented detection systems, we compare their classification performance with several state-of-the-art methods for fake review detection. The proposed systems perform well on all datasets, irrespective of their sentiment polarity and product category.

This is a preview of subscription content, log in via an institution to check access.

Access this article

Log in via an institution

Price excludes VAT (USA)
Tax calculation will be finalised during checkout.

Instant access to the full article PDF.

Institutional subscriptions

Fake news, disinformation and misinformation in social media: a review

Article 09 February 2023

A review on sentiment analysis and emotion detection from text

Article 28 August 2021

Sentiment Analysis in Social Media Data for Depression Detection Using Artificial Intelligence: A Review

Article 19 November 2021

Notes

References

Ahmed K, El Tazi N, Hossny AH (2015) Sentiment analysis over social networks: an overview. In: 2015 IEEE international conference on systems, man, and cybernetics, IEEE, pp 2174–2179. https://doi.org/10.1109/smc.2015.380
Ahmed H, Traore I, Saad S (2018) Detecting opinion spams and fake news using text classification. Secur Priv 1(1):e9. https://doi.org/10.1002/spy2.9
Article Google Scholar
Asghar MZ, Ullah A, Ahmad S, Khan A (2019) Opinion spam detection framework using hybrid classification scheme. Soft Comput. https://doi.org/10.1007/s00500-019-04107-y
Article Google Scholar
Baccianella S, Esuli A, Sebastiani F (2010) Sentiwordnet 3.0: an enhanced lexical resource for sentiment analysis and opinion mining. Lang Resour Eval 10:2200–2204
Google Scholar
Barbado R, Araque O, Iglesias CA (2019) A framework for fake review detection in online consumer electronics retailers. Inf Process Manag 56(4):1234–1244. https://doi.org/10.1016/j.indmarman.2019.08.003
Article Google Scholar
Barushka A, Hajek P (2016) Spam filtering using regularized neural networks with rectified linear units. In: Adorni G, Cagnoni S, Gori M, Maratea M (eds) Conference of the Italian association for artificial intelligence, vol 10037. Lecture notes in computer science. Springer, Cham, pp 65–75. https://doi.org/10.1007/978-3-319-49130-1_6
Chapter Google Scholar
Barushka A, Hajek P (2018) Spam filtering in social networks using regularized deep neural networks with ensemble learning. In: Iliadis L, Maglogiannis I, Plagianakos V (eds) Artificial intelligence applications and innovations. AIAI 2018, vol 519. IFIP advances in information and communication technology. Springer, Cham, pp 38–49. https://doi.org/10.1007/978-3-319-92007-8_4
Chapter Google Scholar
Barushka A, Hajek P (2018) Spam filtering using integrated distribution-based balancing approach and regularized deep neural networks. Appl Intell 48(10):3538–3556. https://doi.org/10.1007/s10489-018-1161-y
Article Google Scholar
Barushka A, Hajek P (2019) Spam detection on social networks using cost-sensitive feature selection and ensemble-based regularized deep neural networks. Neural Comput Appl. https://doi.org/10.1007/s00521-019-04331-5
Article Google Scholar
Barushka A, Hajek P (2019) Review spam detection using word embeddings and deep neural networks. In: MacIntyre J, Maglogiannis I, Iliadis L, Pimenidis E (eds) Artificial intelligence applications and innovations. AIAI 2019, vol 559. IFIP advances in information and communication technology. Springer, Cham, pp 340–350. https://doi.org/10.1007/978-3-030-19823-7_28
Chapter Google Scholar
Bravo-Marquez F, Mendoza M, Poblete B (2014) Meta-level sentiment models for big social data analysis. Knowl Based Syst 69:86–99. https://doi.org/10.1016/j.knosys.2014.05.016
Article Google Scholar
Bravo-Marquez F, Frank E, Mohammad SM, Pfahringer B (2016) Determining word-emotion associations from tweets by multi-label classification. In: 2016 IEEE/WIC/ACM international conference on web intelligence (WI), IEEE, pp 536–539. https://doi.org/10.1109/wi.2016.0091
Brazdil PB, Soares C, Da Costa JP (2003) Ranking learning algorithms: using IBL and meta-learning on accuracy and time results. Mach Learn 50(3):251–277. https://doi.org/10.1023/A:102171390
Article MATH Google Scholar
BrightLocal (2018) Local consumer review survey 2018. https://www.brightlocal.com/research/local-consumer-review-survey/. Accessed 8 Nov 2019
Chandy R, Gu H (2012) Identifying spam in the iOS app store. In: Proceedings of the 2nd joint WICOW/AIRWeb workshop on web quality, ACM, pp 56–59. https://doi.org/10.1145/2184305.2184317
Chatzakou D, Vakali A (2015) Harvesting opinions and emotions from social media textual resources. IEEE Internet Comput 19(4):46–50. https://doi.org/10.1109/MIC.2015.28
Article Google Scholar
Chen W, Yeo CK, Lau CT, Lee BS (2017) A study on real-time low-quality content detection on Twitter from the users’ perspective. PLoS ONE 12(8):e0182487. https://doi.org/10.1371/journal.pone.0182487
Article Google Scholar
Crawford M, Khoshgoftaar TM, Prusa JD, Richter AN, Al Najada H (2015) Survey of review spam detection using machine learning techniques. J Big Data 2(1):1–23. https://doi.org/10.1186/s40537-015-0029-9
Article Google Scholar
Elmurngi E, Gherbi A (2017) An empirical study on detecting fake reviews using machine learning techniques. In: 7th international conference on innovative computing technology (INTECH), IEEE, pp 107–114. https://doi.org/10.1109/intech.2017.8102442
Felbermayr A, Nanopoulos A (2016) The role of emotions for the perceived usefulness in online customer reviews. J Interact Mark 36:60–76. https://doi.org/10.1016/j.intmar.2016.05.004
Article Google Scholar
Floyd K, Freling R, Alhoqail S, Cho HY, Freling T (2014) How online product reviews affect retail sales: a meta-analysis. J Retail 90(2):217–232. https://doi.org/10.1016/j.jretai.2014.04.004
Article Google Scholar
Garcia L (2018) Deception on Amazon—an NLP exploration. https://medium.com/@lievgarcia/deception-on-amazon-c1e30d977cfd. Accessed 01 Sept 2019
Garcia S, Fernandez A, Luengo J, Herrera F (2010) Advanced nonparametric tests for multiple comparisons in the design of experiments in computational intelligence and data mining: experimental analysis of power. Inf Sci 180(10):2044–2064. https://doi.org/10.1016/j.ins.2009.12.010
Article Google Scholar
Ghai R, Kumar S, Pandey AC (2019) Spam detection using rating and review processing method. In: Panigrahi B, Trivedi M, Mishra K, Tiwari S, Singh P (eds) Smart innovations in communication and computational sciences. Springer, Singapore, pp 189–198. https://doi.org/10.1007/978-981-10-8971-8_18
Chapter Google Scholar
Hajek P (2018) Combining bag-of-words and sentiment features of annual reports to predict abnormal stock returns. Neural Comput Appl 29(7):343–358. https://doi.org/10.1007/s00521-017-3194-2
Article Google Scholar
Harris C (2012) Detecting deceptive opinion spam using human computation. In: Workshops at AAAI on artificial intelligence, AAAI, pp 87–93
He R, McAuley J (2016) Ups and downs: modeling the visual evolution of fashion trends with one-class collaborative filtering. In: Proceedings of the 25th international conference on world wide web, pp 507–517. https://doi.org/10.1145/2872427.2883037
Heydari A, Ali Tavakoli M, Salim N, Heydari Z (2015) Detection of review spam: a survey. Expert Syst Appl 42(7):3634–3642. https://doi.org/10.1016/j.eswa.2014.12.029
Article Google Scholar
Hu M, Liu B (2004) Mining and summarizing customer reviews. In: Proceedings of the 10th ACM SIGKDD international conference on knowledge discovery and data mining, ACM, pp 168–177. https://doi.org/10.1145/1014052.1014073
Hussain N, Turab Mirza H, Rasool G, Hussain I, Kaleem M (2019) Spam review detection techniques: a systematic literature review. Appl Sci 9(5):987. https://doi.org/10.3390/app9050987
Article Google Scholar
Ikeda K, Hattori G, Ono C, Asoh H, Higashino T (2013) Twitter user profiling based on text and community mining for market analysis. Knowl Based Syst 51:35–47. https://doi.org/10.1016/j.knosys.2013.06.020
Article Google Scholar
Jain G, Sharma M, Agarwal B (2018) Spam detection on social media using semantic convolutional neural network. Int J Knowl Discov Bioinform (IJKDB) 8(1):12–26. https://doi.org/10.4018/IJKDB.2018010102
Article Google Scholar
Jain G, Sharma M, Agarwal B (2019) Spam detection in social media using convolutional and long short term memory neural network. Ann Math Artif Intell 85(1):21–44. https://doi.org/10.1007/s10472-018-9612-z
Article MATH Google Scholar
Jindal N, Liu B (2007) Analyzing and detecting review spam. In: 7th IEEE international conference on data mining, ICDM 2007, IEEE, pp 547–552. https://doi.org/10.1109/icdm.2007.68
Kennedy S, Walsh N, Sloka K, McCarren A, Foster J (2019) Fact or factitious? Contextualized opinion spam detection. In: Proceedings of the 57th annual meeting of the association for computational linguistics: student research workshop, ACL, pp 344–350. https://doi.org/10.18653/v1/p19-2048
Kiritchenko S, Zhu X, Mohammad SM (2014) Sentiment analysis of short informal texts. J Artif Intell Res 50:723–762
Article Google Scholar
Lau RY, Liao SY, Kwok RCW, Xu K, Xia Y, Li Y (2011) Text mining and probabilistic language modeling for online review spam detecting. ACM Trans Manag Inf Syst 2(4):1–30. https://doi.org/10.1145/2070710.2070716
Article Google Scholar
Le Q, Mikolov T (2014) Distributed representations of sentences and documents. In: International conference on machine learning, JMLR, vol 32, pp 1188–1196
Li F, Huang M, Yang Y, Zhu X (2011) Learning to identify review spam. In: International joint conference on artificial intelligence (IJCAI 2011), pp 2488–2493
Li J, Ott M, Cardie C, Hovy E (2014) Towards a general rule for identifying deceptive opinion spam. In: Proceedings of the 52nd annual meeting of the association for computational linguistics, ACL, vol 1, pp 1566–1576. https://doi.org/10.3115/v1/p14-1147
Li H, Chen Z, Mukherjee A, Liu B, Shao J (2015) Analyzing and detecting opinion spam on a large-scale dataset via temporal and spatial patterns. In: 9th international AAAI conference on web and social media (ICWSM 2015), AAAI, pp 634–637
Li H, Fei G, Wang S, Liu B, Shao W, Mukherjee A, Shao J (2017) Bimodal distribution and co-bursting in review spam detection. In: 26th international conference on world wide web, ACM, pp 1063–1072. https://doi.org/10.1145/3038912.3052582
Li L, Qin B, Ren W, Liu T (2017) Document representation and feature combination for deceptive spam review detection. Neurocomputing 254:33–41. https://doi.org/10.1016/j.neucom.2016.10.080
Article Google Scholar
Lim EP, Nguyen VA, Jindal N, Liu B, Lauw HW (2010) Detecting product review spammers using rating behaviors. In: 19th ACM international conference on information and knowledge management, ACM, pp 939–948. https://doi.org/10.1145/1871437.1871557
Lin M, Chen Q, Yan S (2014) Network in network. In: International conference on learning representations (ICLR), ICLR, pp 1–10
Liu Y, Pang B (2018) A unified framework for detecting author spamicity by modeling review deviation. Expert Syst Appl 112:148–155. https://doi.org/10.1016/j.eswa.2018.06.028
Article Google Scholar
Liu Y, Pang B, Wang X (2019) Opinion spam detection by incorporating multimodal embedded representation into a probabilistic review graph. Neurocomputing 366:276–283. https://doi.org/10.1016/j.neucom.2019.08.013
Article Google Scholar
Madisetty S, Desarkar MS (2018) A neural network-based ensemble approach for spam detection in Twitter. IEEE Trans Comput Soc Syst 5(4):973–984. https://doi.org/10.1109/TCSS.2018.2878852
Article Google Scholar
Malik MSI, Hussain A (2017) Helpfulness of product reviews as a function of discrete positive and negative emotions. Comput Hum Behav 73:290–302. https://doi.org/10.1016/j.chb.2017.03.053
Article Google Scholar
Mikolov T, Sutskever I, Chen K, Corrado GS, Dean J (2013) Distributed representations of words and phrases and their compositionality. In: Advances in neural information processing systems, NIPS, vol 26, pp 3111–3119
Mohammad SM, Turney PD (2013) Crowdsourcing a word–emotion association lexicon. Comput Intell 29(3):436–465. https://doi.org/10.1111/j.1467-8640.2012.00460.x
Article MathSciNet Google Scholar
Mukherjee A, Venkataraman V, Liu B, Glance N (2013) What yelp fake review filter might be doing?. In: 7th international AAAI conference on weblogs and social media, AAAI, pp 409–418
Nielsen FÅ (2011) A new ANEW: evaluation of a word list for sentiment analysis in microblogs. In: Proceedings of the ESWC2011 workshop on ‘making sense of microposts’: big things come in small packages, pp 93–98
Ott M, Choi Y, Cardie C, Hancock JT (2011) Finding deceptive opinion spam by any stretch of the imagination. In: Proceedings of the 49th annual meeting of the association for computational linguistics: human language technologies, ACL, vol 1, pp 309–319
Ott M, Cardie C, Hancock J (2012) Estimating the prevalence of deception in online review communities. In: 21st international conference on world wide web, ACM, pp 201–210. https://doi.org/10.1145/2187836.2187864
Ott M, Cardie C, Hancock JT (2013) Negative deceptive opinion spam. In: 2013 conference of the North American chapter of the association for computational linguistics: human language technologies, ACL, pp 497–501
Pandey AC, Rajpoot DS (2019) Spam review detection using spiral cuckoo search clustering method. Evol Intell 12(2):147–164. https://doi.org/10.1007/s12065-019-00204-x
Article Google Scholar
Patel NA, Patel R (2018) A survey on fake review detection using machine learning techniques. In: 2018 4th international conference on computing communication and automation (ICCCA), IEEE, pp 1–6. https://doi.org/10.1109/ccaa.2018.8777594
Peng Q, Zhong M (2014) Detecting spam review through sentiment analysis. J Softw 9(8):2065–2072. https://doi.org/10.4304/jsw.9.8.2065-2072
Article Google Scholar
Rayana S, Akoglu L (2015) Collective opinion spam detection: bridging review networks and metadata. In: 21th ACM SIGKDD international conference on knowledge discovery and data mining, ACM, pp 985–994. https://doi.org/10.1145/2783258.2783370
Ren Y, Ji D (2017) Neural networks for deceptive opinion spam detection: an empirical study. Inf Sci 385:213–224. https://doi.org/10.1016/j.ins.2017.01.015
Article Google Scholar
Rout JK, Dalmia A, Choo KKR, Bakshi S, Jena SK (2017) Revisiting semi-supervised learning for online deceptive review detection. IEEE Access 5:1319–1327. https://doi.org/10.1109/ACCESS.2017.2655032
Article Google Scholar
Rout JK, Dash AK, Ray NK (2018) A framework for fake review detection: issues and challenges. In: 2018 international conference on information technology (ICIT), IEEE, pp 7–10. https://doi.org/10.1109/icit.2018.00014
Shojaee S, Murad MAA, Azman AB, Sharef NM, Nadali S (2013) Detecting deceptive reviews using lexical and syntactic features. In: 13th international conference on intelligent systems design and applications, IEEE, pp 53–58. https://doi.org/10.1109/isda.2013.6920707
Sun C, Du Q, Tian G (2016) Exploiting product related review features for fake review detection. Math Probl Eng 2016:1–7. https://doi.org/10.1155/2016/4935792
Article Google Scholar
Tang X, Qian T, You Z (2019) Generating behavior features for cold-start spam review detection. In: International conference on database systems for advanced applications, Springer, Cham, pp 324–328. https://doi.org/10.1007/978-3-030-18590-9_38
The Times (2018) ‘A third of TripAdvisor reviews are fake’ as cheats buy five stars. The Times September 22, 2018. https://www.thetimes.co.uk/article/hotel-and-caf-cheats-are-caught-trying-to-buy-tripadvisor-stars-027fbcwc8. Accessed 22 Jan 2019
TripAdvisor Homepage. http://ir.tripadvisor.com/. Accessed 21 Jan 2019
Vidanagama DU, Silva TP, Karunananda AS (2019) Deceptive consumer review detection: a survey. Artif Intell Rev. https://doi.org/10.1007/s10462-019-09697-5
Article Google Scholar
Wang G, Xie S, Liu B, Philip SY (2011) Review graph based online store review spammer detection. In: 11th international conference on data mining (ICDM 2011), IEEE, pp 1242–1247. https://doi.org/10.1109/icdm.2011.124
Wang G, Li C, Wang W, Zhang Y, Shen D, Zhang X, Henao R, Carin L (2018) Joint embedding of words and labels for text classification. In: Proceedings of the 56th annual meeting of the association for computational linguistics, ACL, pp 2321–2331. https://doi.org/10.18653/v1/p18-1216
Wilson T, Wiebe J, Hoffmann P (2005) Recognizing contextual polarity in phrase-level sentiment analysis. In: Proceedings of human language technology conference and conference on empirical methods in natural language processing, ACL, pp 347–354
Xie S, Wang G, Lin S, Yu PS (2012) Review spam detection via temporal pattern discovery. In: 18th ACM SIGKDD international conference on knowledge discovery and data mining, ACM, pp 823–831. https://doi.org/10.1145/2339530.2339662
Xue H, Wang Q, Luo B, Seo H, Li F (2019) Content-aware trust propagation toward online review spam detection. J Data Inf Qual (JDIQ) 11(3):11. https://doi.org/10.1145/3305258
Article Google Scholar
Ye J, Kumar S, Akoglu L (2016) Temporal opinion spam detection by multivariate indicative signals. In: 10th international AAAI conference on web and social media (ICWSM 2016), AAAI, pp 743–746
Yilmaz CM, Durahim AO (2018) SPR2EP: a semi-supervised spam review detection framework. In: 2018 IEEE/ACM international conference on advances in social networks analysis and mining (ASONAM), IEEE, pp 306–313. https://doi.org/10.1109/asonam.2018.8508314
Zeng ZY, Lin JJ, Chen MS, Chen MH, Lan YQ, Liu JL (2019) A review structure based ensemble model for deceptive review spam. Information 10(7):243. https://doi.org/10.3390/info10070243
Article Google Scholar

Download references

Acknowledgements

This article was supported by the scientific research project of the Czech Sciences Foundation Grant No: 19-15498S and by the Operational Program: Research and Innovation project “Fake news on the Internet—identification, content analysis, emotions”, co-funded by the European Regional Development Fund.

Author information

Authors and Affiliations

Faculty of Economics and Administration, Institute of System Engineering and Informatics, University of Pardubice, Studentská 84, 532 10, Pardubice, Czech Republic
Petr Hajek & Aliaksandr Barushka
Department of Computer Science, Constantine the Philosopher University in Nitra, 949 74, Nitra, Slovakia
Michal Munk

Authors

Petr Hajek
View author publications
You can also search for this author in PubMed Google Scholar
Aliaksandr Barushka
View author publications
You can also search for this author in PubMed Google Scholar
Michal Munk
View author publications
You can also search for this author in PubMed Google Scholar

Corresponding author

Correspondence to Petr Hajek.

Ethics declarations

Conflict of interest

The authors declare that they have no conflict of interest.

Additional information

Publisher's Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Rights and permissions

Reprints and permissions

About this article

Cite this article

Hajek, P., Barushka, A. & Munk, M. Fake consumer review detection using deep neural networks integrating word embeddings and emotion mining. Neural Comput & Applic 32, 17259–17274 (2020). https://doi.org/10.1007/s00521-020-04757-2

Download citation

Received: 09 November 2019
Accepted: 23 January 2020
Published: 01 February 2020
Issue Date: December 2020
DOI: https://doi.org/10.1007/s00521-020-04757-2

Keywords

Access this article

Log in via an institution

Price excludes VAT (USA)
Tax calculation will be finalised during checkout.

Instant access to the full article PDF.

Institutional subscriptions

Fake consumer review detection using deep neural networks integrating word embeddings and emotion mining

Abstract

Access this article

Similar content being viewed by others

Fake news, disinformation and misinformation in social media: a review

A review on sentiment analysis and emotion detection from text

Sentiment Analysis in Social Media Data for Depression Detection Using Artificial Intelligence: A Review

Notes

References

Acknowledgements

Author information

Authors and Affiliations

Corresponding author

Ethics declarations

Conflict of interest

Additional information

Publisher's Note

Rights and permissions

About this article

Cite this article

Keywords

Navigation

Fake consumer review detection using deep neural networks integrating word embeddings and emotion mining

Abstract

Access this article

Similar content being viewed by others

Fake news, disinformation and misinformation in social media: a review

A review on sentiment analysis and emotion detection from text

Sentiment Analysis in Social Media Data for Depression Detection Using Artificial Intelligence: A Review

Notes

References

Acknowledgements

Author information

Authors and Affiliations

Corresponding author

Ethics declarations

Conflict of interest

Additional information

Publisher's Note

Rights and permissions

About this article

Cite this article

Share this article

Keywords

Search

Navigation