Arrow Research search

Author name cluster

Sneha Mehta

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

3 papers
1 author row

Possible papers

3

NeurIPS Conference 2022 Conference Paper

TweetNERD - End to End Entity Linking Benchmark for Tweets

  • Shubhanshu Mishra
  • Aman Saini
  • Raheleh Makki
  • Sneha Mehta
  • Aria Haghighi
  • Ali Mollahosseini

Named Entity Recognition and Disambiguation (NERD) systems are foundational for information retrieval, question answering, event detection, and other natural language processing (NLP) applications. We introduce TweetNERD, a dataset of 340K+ Tweets across 2010-2021, for benchmarking NERD systems on Tweets. This is the largest and most temporally diverse open sourced dataset benchmark for NERD on Tweets and can be used to facilitate research in this area. We describe evaluation setup with TweetNERD for three NERD tasks: Named Entity Recognition (NER), Entity Linking with True Spans (EL), and End to End Entity Linking (End2End); and provide performance of existing publicly available methods on specific TweetNERD splits. TweetNERD is available at: https: //doi. org/10. 5281/zenodo. 6617192 under Creative Commons Attribution 4. 0 International (CC BY 4. 0) license. Check out more details at https: //github. com/twitter-research/TweetNERD.

AAAI Conference 2020 Conference Paper

Simplify-Then-Translate: Automatic Preprocessing for Black-Box Translation

  • Sneha Mehta
  • Bahareh Azarnoush
  • Boris Chen
  • Avneesh Saluja
  • Vinith Misra
  • Ballav Bihani
  • Ritwik Kumar

Black-box machine translation systems have proven incredibly useful for a variety of applications yet by design are hard to adapt, tune to a specific domain, or build on top of. In this work, we introduce a method to improve such systems via automatic pre-processing (APP) using sentence simplification. We first propose a method to automatically generate a large in-domain paraphrase corpus through back-translation with a black-box MT system, which is used to train a paraphrase model that “simplifies” the original sentence to be more conducive for translation. The model is used to preprocess source sentences of multiple low-resource language pairs. We show that this preprocessing leads to better translation performance as compared to non-preprocessed source sentences. We further perform side-by-side human evaluation to verify that translations of the simplified sentences are better than the original ones. Finally, we provide some guidance on recommended language pairs for generating the simplification model corpora by investigating the relationship between ease of translation of a language pair (as measured by BLEU) and quality of the resulting simplification model from backtranslations of this language pair (as measured by SARI), and tie this into the downstream task of low-resource translation.

IJCAI Conference 2020 Conference Paper

Supporting Historical Photo Identification with Face Recognition and Crowdsourced Human Expertise (Extended Abstract)

  • Vikram Mohanty
  • David Thames
  • Sneha Mehta
  • Kurt Luther

Identifying people in historical photographs is important for interpreting material culture, correcting the historical record, and creating economic value, but it is also a complex and challenging task. In this paper, we focus on identifying portraits of soldiers who participated in the American Civil War (1861-65). Millions of these portraits survive, but only 10-20% are identified. We created Photo Sleuth, a web-based platform that combines crowdsourced human expertise and automated face recognition to support Civil War portrait identification. Our mixed-methods evaluation of Photo Sleuth one month after its public launch showed that it helped users successfully identify unknown portraits.

v2026.09.13