Terpene

Authors
Sanjaya Wijeratne, Lakshika Balasuriya, Amit Sheth, Derek Doran
Publication date
2017/5/3
Journal
Proceedings of the International AAAI Conference on Web and Social Media
Volume
11
Issue
1
Pages
437-446
Description
This paper presents the release of EmojiNet, the largest machine-readable emoji sense inventory that links Unicode emoji representations to their English meanings extracted from the Web. EmojiNet is a dataset consisting of:(i) 12,904 sense labels over 2,389 emoji, which were extracted from the web and linked to machine-readable sense definitions seen in BabelNet;(ii) context words associated with each emoji sense, which are inferred through word embedding models trained over Google News corpus and a Twitter message corpus for each emoji sense definition; and (iii) recognizing discrepancies in the presentation of emoji on different platforms, specification of the most likely platform-based emoji sense for a selected set of emoji. The dataset is hosted as an open service with a REST API and is available at http://emojinet. knoesis. org/. The development of this dataset, evaluation of its quality, and its applications including emoji sense disambiguation and emoji sense similarity are discussed.
Total citations
2017201820192020202120222023202442491781645
Scholar articles
S Wijeratne, L Balasuriya, A Sheth, D Doran - Proceedings of the International AAAI Conference on …, 2017

Leave a Reply