Publications
Sort & Filters
Filters
22 results
22 results
Working Paper
Carlson, Jacob, Tom Bryan, and Melissa Dell. “Efficient OCR for Building a Diverse Digital History.”
Carlson, Jacob, Tom Bryan, and Melissa Dell. “Efficient OCR for Building a Diverse Digital History.”
Thousands of users consult digital archives daily, but the information they can access is unrepresentative of the diversity of documentary history. The sequence-to-sequence architecture typically used for optical character recognition (OCR) – which...
Arora, Abhishek, Xinmei Yang, Shao Yu Jheng, and Melissa Dell. “Linking Representations With Multimodal Contrastive Learning.”
Arora, Abhishek, Xinmei Yang, Shao Yu Jheng, and Melissa Dell. “Linking Representations With Multimodal Contrastive Learning.”
Many applications require grouping instances contained in diverse document datasets into classes. Most widely used methods do not employ deep learning and do not exploit the inherently multimodal nature of documents. Notably, record linkage is typically...
Yang, Xinmei, Abhishek Arora, Shao-Yu Jheng, and Melissa Dell. “Quantifying Character Similarity With Vision Transformers.”
Yang, Xinmei, Abhishek Arora, Shao-Yu Jheng, and Melissa Dell. “Quantifying Character Similarity With Vision Transformers.”
Record linkage is a bedrock of quantitative social science, as analyses often require linking data from multiple, noisy sources. Off-the-shelf string matching methods are widely used, as they are straightforward and cheap to implement and scale. Not all...
Silcock, Emily, and Melissa Dell. “A Massive Scale Semantic Similarity Dataset of Historical English.”
Silcock, Emily, and Melissa Dell. “A Massive Scale Semantic Similarity Dataset of Historical English.”
A diversity of tasks use language models trained on semantic similarity data. While there are a variety of datasets that capture semantic similarity, they are either constructed from modern web data or are relatively small datasets created in the past...
Dell, Melissa, Jacob Carlson, Tom Bryan, Emily Silcock, Abhishek Arora, Zejiang Shen, Luca D’Amico-Wong, Quan Le, Pablo Querubin, and Leander Heldring. “American Stories: A Large-Scale Structured Text Dataset of Historical U.S. Newspapers.”
Dell, Melissa, Jacob Carlson, Tom Bryan, Emily Silcock, Abhishek Arora, Zejiang Shen, Luca D’Amico-Wong, Quan Le, Pablo Querubin, and Leander Heldring. “American Stories: A Large-Scale Structured Text Dataset of Historical U.S. Newspapers.”
Existing full text datasets of U.S. public domain newspapers do not recognize the often complex layouts of newspaper scans, and as a result the digitized content scrambles texts from articles, headlines, captions, advertisements, and other layout regions...
Arora, Abhishek, and Melissa Dell. “LinkTransformer: A Unified Package for Record Linkage With Transformer Language Models.”
Arora, Abhishek, and Melissa Dell. “LinkTransformer: A Unified Package for Record Linkage With Transformer Language Models.”
Linking information across sources is fundamental to a variety of analyses in social science, business, and government. While large language models (LLMs) offer enormous promise for improving record linkage in noisy datasets, in many domains approximate...
Forthcoming
Shen, Zejiang, Jian Zhao, Weining Li, Yaoliang Yu, and Melissa Dell. “OLALA: Object-Level Active Learning Based Layout Annotation”. EMNLP Computational Social Science Workshop, n.d., 2023.
Shen, Zejiang, Jian Zhao, Weining Li, Yaoliang Yu, and Melissa Dell. “OLALA: Object-Level Active Learning Based Layout Annotation”. EMNLP Computational Social Science Workshop, n.d., 2023.
Layout detection is an essential step for accurately extracting structured contents from historical documents. The intricate and varied layouts present in these document images make it expensive to label the numerous layout regions that can be densely...
Silcock, Emily, Luca D’Amico-Wong, Jinglin Yang, and Melissa Dell. “Noise-Robust De-Duplication at Scale”. International Conference on Learning Representations, n.d.
Silcock, Emily, Luca D’Amico-Wong, Jinglin Yang, and Melissa Dell. “Noise-Robust De-Duplication at Scale”. International Conference on Learning Representations, n.d.
Identifying near duplicates within large, noisy text corpora has a myriad of applications that range from de-duplicating training datasets, reducing privacy risk, and evaluating test set leakage, to identifying reproduced news articles and literature...
2021
Shen, Zejiang, Ruochen Zhang, Melissa Dell, Benjamin Lee, Jacob Carlson, and Weining Li. “LayoutParser: A Unified Toolkit for Deep Learning Based Document Image Analysis”. International Conference on Document Analysis and Recognition, 2021, 131-46.
Shen, Zejiang, Ruochen Zhang, Melissa Dell, Benjamin Lee, Jacob Carlson, and Weining Li. “LayoutParser: A Unified Toolkit for Deep Learning Based Document Image Analysis”. International Conference on Document Analysis and Recognition, 2021, 131-46.
Recent advances in document image analysis (DIA) have been primarily driven by the application of neural networks. Ideally, research outcomes could be easily deployed in production and extended for further investigation. However, various factors like...
2020
Dell, Melissa, and Benjamin Olken. “The Development Effects of the Extractive Colonial Economy: The Dutch Cultivation System in Java”. Review of Economic Studies 87, no. 1 (2020): 164-203.
Dell, Melissa, and Benjamin Olken. “The Development Effects of the Extractive Colonial Economy: The Dutch Cultivation System in Java”. Review of Economic Studies 87, no. 1 (2020): 164-203.
Shen, Zejiang, Kaixuan Zhang, and Melissa Dell. “A Large Dataset of Historical Japanese Documents With Complex Layouts”. IEEE/CVF/Conference/on/Computer/Vision/and/Pattern/Recognition/Workshops, 2020, 548-59.
Shen, Zejiang, Kaixuan Zhang, and Melissa Dell. “A Large Dataset of Historical Japanese Documents With Complex Layouts”. IEEE/CVF/Conference/on/Computer/Vision/and/Pattern/Recognition/Workshops, 2020, 548-59.
Deep learning-based approaches for automatic document layout analysis and content extraction have the potential to unlock rich information trapped in historical documents on a large scale. One major hurdle is the lack of large datasets for training...