#  Publications 

 



 



 Sort &amp; Filters  close  

## Filters

    Publication type expand\_more    Journal Article (15) 

  Working Paper (7) 



 

  



 

  Search Within Results  

  Search Within Results search  

##  22 results 

  Show filters filter\_alt    Sort by Year of PublicationAlphabetical A-Z sort



 

##  22 results 

  Download 22 citations  download- [BibTeX](/node/707571/export?format=bibtex)
- [EndNote X3 XML](/node/707571/export?format=endnote8)
- [EndNote 7 XML](/node/707571/export?format=endnote7)
- [Endnote tagged](/node/707571/export?format=tagged)
- [Marc](/node/707571/export?format=marc)
- [PubMedId](/node/707571/export?format=pubmed_id)
- [RIS](/node/707571/export?format=ris)
 


 

### Working Paper

Dell, M. “[Path Dependence in Development: Evidence from the Mexican Revolution](/publications/path-dependence-development-evidence-mexican-revolution).”



 

 

Dell, M. “[Path Dependence in Development: Evidence from the Mexican Revolution](/publications/path-dependence-development-evidence-mexican-revolution).”



 

 

 

- [ picture\_as\_pdfPDF](/sites/g/files/omnuum7696/files/dell/files/revolutiondraft.pdf)
- [ picture\_as\_pdfData appendix](/sites/g/files/omnuum7696/files/dell/files/data_construction_appendix_mxrev.pdf)
 
- [ picture\_as\_pdfPDF](/sites/g/files/omnuum7696/files/dell/files/revolutiondraft.pdf)
- [ picture\_as\_pdfData appendix](/sites/g/files/omnuum7696/files/dell/files/data_construction_appendix_mxrev.pdf)
 
 

Carlson, Jacob, Tom Bryan, and Melissa Dell. “[Efficient OCR for Building a Diverse Digital History](/publications/efficient-ocr-building-diverse-digital-history).”



 

 

Carlson, Jacob, Tom Bryan, and Melissa Dell. “[Efficient OCR for Building a Diverse Digital History](/publications/efficient-ocr-building-diverse-digital-history).”



 

 

 

- add\_circle do\_not\_disturb\_on Abstract
- [ picture\_as\_pdfPaper](/sites/g/files/omnuum7696/files/dell/files/effocr.pdf)
 
 Thousands of users consult digital archives daily, but the information they can access is unrepresentative of the diversity of documentary history. The sequence-to-sequence architecture typically used for optical character recognition (OCR) – which... 

 

 

- [ picture\_as\_pdfPaper](/sites/g/files/omnuum7696/files/dell/files/effocr.pdf)
 
 

Arora, Abhishek, Xinmei Yang, Shao Yu Jheng, and Melissa Dell. “[Linking Representations With Multimodal Contrastive Learning](/publications/linking-representations-multimodal-contrastive-learning).”



 

 

Arora, Abhishek, Xinmei Yang, Shao Yu Jheng, and Melissa Dell. “[Linking Representations With Multimodal Contrastive Learning](/publications/linking-representations-multimodal-contrastive-learning).”



 

 

 

- add\_circle do\_not\_disturb\_on Abstract
- [ picture\_as\_pdfPaper](/sites/g/files/omnuum7696/files/dell/files/rl.pdf)
 
 Many applications require grouping instances contained in diverse document datasets into classes. Most widely used methods do not employ deep learning and do not exploit the inherently multimodal nature of documents. Notably, record linkage is typically...



 

 

- [ picture\_as\_pdfPaper](/sites/g/files/omnuum7696/files/dell/files/rl.pdf)
 
 

Yang, Xinmei, Abhishek Arora, Shao-Yu Jheng, and Melissa Dell. “[Quantifying Character Similarity With Vision Transformers](/publications/quantifying-character-similarity-vision-transformers).”



 

 

Yang, Xinmei, Abhishek Arora, Shao-Yu Jheng, and Melissa Dell. “[Quantifying Character Similarity With Vision Transformers](/publications/quantifying-character-similarity-vision-transformers).”



 

 

 

- add\_circle do\_not\_disturb\_on Abstract
- [ picture\_as\_pdfPaper](/sites/g/files/omnuum7696/files/dell/files/homoglyphs.pdf)
 
 Record linkage is a bedrock of quantitative social science, as analyses often require linking data from multiple, noisy sources. Off-the-shelf string matching methods are widely used, as they are straightforward and cheap to implement and scale. Not all... 

 

 

- [ picture\_as\_pdfPaper](/sites/g/files/omnuum7696/files/dell/files/homoglyphs.pdf)
 
 

Silcock, Emily, and Melissa Dell. “[A Massive Scale Semantic Similarity Dataset of Historical English](/publications/massive-scale-semantic-similarity-dataset-historical-english).”



 

 

Silcock, Emily, and Melissa Dell. “[A Massive Scale Semantic Similarity Dataset of Historical English](/publications/massive-scale-semantic-similarity-dataset-historical-english).”



 

 

 

- add\_circle do\_not\_disturb\_on Abstract
- [ picture\_as\_pdfPaper](/sites/g/files/omnuum7696/files/dell/files/qshfjvnkkknzwdfrbkvvvmmbygtmtymn.pdf)
 
 A diversity of tasks use language models trained on semantic similarity data. While there are a variety of datasets that capture semantic similarity, they are either constructed from modern web data or are relatively small datasets created in the past... 

 

 

- [ picture\_as\_pdfPaper](/sites/g/files/omnuum7696/files/dell/files/qshfjvnkkknzwdfrbkvvvmmbygtmtymn.pdf)
 
 

Dell, Melissa, Jacob Carlson, Tom Bryan, Emily Silcock, Abhishek Arora, Zejiang Shen, Luca D’Amico-Wong, Quan Le, Pablo Querubin, and Leander Heldring. “[American Stories: A Large-Scale Structured Text Dataset of Historical U.S. Newspapers](/publications/american-stories-large-scale-structured-text-dataset-historical-us-newspapers).”



 

 

Dell, Melissa, Jacob Carlson, Tom Bryan, Emily Silcock, Abhishek Arora, Zejiang Shen, Luca D’Amico-Wong, Quan Le, Pablo Querubin, and Leander Heldring. “[American Stories: A Large-Scale Structured Text Dataset of Historical U.S. Newspapers](/publications/american-stories-large-scale-structured-text-dataset-historical-us-newspapers).”



 

 

 

- add\_circle do\_not\_disturb\_on Abstract
- [ picture\_as\_pdfPaper](/sites/g/files/omnuum7696/files/dell/files/americanstories.pdf)
 
 Existing full text datasets of U.S. public domain newspapers do not recognize the often complex layouts of newspaper scans, and as a result the digitized content scrambles texts from articles, headlines, captions, advertisements, and other layout regions... 

 

 

- [ picture\_as\_pdfPaper](/sites/g/files/omnuum7696/files/dell/files/americanstories.pdf)
 
 

Arora, Abhishek, and Melissa Dell. “[LinkTransformer: A Unified Package for Record Linkage With Transformer Language Models](/publications/linktransformer-unified-package-record-linkage-transformer-language-models).”



 

 

Arora, Abhishek, and Melissa Dell. “[LinkTransformer: A Unified Package for Record Linkage With Transformer Language Models](/publications/linktransformer-unified-package-record-linkage-transformer-language-models).”



 

 

 

- add\_circle do\_not\_disturb\_on Abstract
- [ picture\_as\_pdfPaper](/sites/g/files/omnuum7696/files/dell/files/linkt.pdf)
 
 Linking information across sources is fundamental to a variety of analyses in social science, business, and government. While large language models (LLMs) offer enormous promise for improving record linkage in noisy datasets, in many domains approximate... 

 

 

- [ picture\_as\_pdfPaper](/sites/g/files/omnuum7696/files/dell/files/linkt.pdf)
 
 

 



### Forthcoming

Shen, Zejiang, Jian Zhao, Weining Li, Yaoliang Yu, and Melissa Dell. “[OLALA: Object-Level Active Learning Based Layout Annotation](/publications/olala-object-level-active-learning-based-layout-annotation)”. *EMNLP Computational Social Science Workshop*, n.d., 2023.



 

 

Shen, Zejiang, Jian Zhao, Weining Li, Yaoliang Yu, and Melissa Dell. “[OLALA: Object-Level Active Learning Based Layout Annotation](/publications/olala-object-level-active-learning-based-layout-annotation)”. *EMNLP Computational Social Science Workshop*, n.d., 2023.



 

 

 

- add\_circle do\_not\_disturb\_on Abstract
- [ picture\_as\_pdfPaper](/sites/g/files/omnuum7696/files/dell/files/nlpcss_1.pdf)
 
 Layout detection is an essential step for accurately extracting structured contents from historical documents. The intricate and varied layouts present in these document images make it expensive to label the numerous layout regions that can be densely... 

 

 

- [ picture\_as\_pdfPaper](/sites/g/files/omnuum7696/files/dell/files/nlpcss_1.pdf)
 
 

Silcock, Emily, Luca D’Amico-Wong, Jinglin Yang, and Melissa Dell. “[Noise-Robust De-Duplication at Scale](/publications/noise-robust-de-duplication-scale)”. *International Conference on Learning Representations*, n.d.



 

 

Silcock, Emily, Luca D’Amico-Wong, Jinglin Yang, and Melissa Dell. “[Noise-Robust De-Duplication at Scale](/publications/noise-robust-de-duplication-scale)”. *International Conference on Learning Representations*, n.d.



 

 

 

- add\_circle do\_not\_disturb\_on Abstract
- [ picture\_as\_pdfPaper](/sites/g/files/omnuum7696/files/dell/files/iclr_samesource_revision_non_anonymous_1.pdf)
 
 Identifying near duplicates within large, noisy text corpora has a myriad of applications that range from de-duplicating training datasets, reducing privacy risk, and evaluating test set leakage, to identifying reproduced news articles and literature... 

 

 

- [ picture\_as\_pdfPaper](/sites/g/files/omnuum7696/files/dell/files/iclr_samesource_revision_non_anonymous_1.pdf)
 
 

 



### 2021

Shen, Zejiang, Ruochen Zhang, Melissa Dell, Benjamin Lee, Jacob Carlson, and Weining Li. “[LayoutParser: A Unified Toolkit for Deep Learning Based Document Image Analysis](/publications/layoutparser-unified-toolkit-deep-learning-based-document-image-analysis)”. *International Conference on Document Analysis and Recognition*, 2021, 131-46.



 

 

Shen, Zejiang, Ruochen Zhang, Melissa Dell, Benjamin Lee, Jacob Carlson, and Weining Li. “[LayoutParser: A Unified Toolkit for Deep Learning Based Document Image Analysis](/publications/layoutparser-unified-toolkit-deep-learning-based-document-image-analysis)”. *International Conference on Document Analysis and Recognition*, 2021, 131-46.



 

 

 

- add\_circle do\_not\_disturb\_on Abstract
- [ descriptionPublisher's Version](https://arxiv.org/pdf/2103.15348.pdf)
 
 Recent advances in document image analysis (DIA) have been primarily driven by the application of neural networks. Ideally, research outcomes could be easily deployed in production and extended for further investigation. However, various factors like... 

 

 

- [ descriptionPublisher's Version](https://arxiv.org/pdf/2103.15348.pdf)
 
 

 



### 2020

Dell, Melissa, and Benjamin Olken. “[The Development Effects of the Extractive Colonial Economy: The Dutch Cultivation System in Java](/publications/development-effects-extractive-colonial-economy-dutch-cultivation-system-java)”. *Review of Economic Studies* 87, no. 1 (2020): 164-203.



 

 

Dell, Melissa, and Benjamin Olken. “[The Development Effects of the Extractive Colonial Economy: The Dutch Cultivation System in Java](/publications/development-effects-extractive-colonial-economy-dutch-cultivation-system-java)”. *Review of Economic Studies* 87, no. 1 (2020): 164-203.



 

 

 

- [ descriptionPublisher's Version](https://www.dropbox.com/sh/xpfzjx5pzfzgktv/AABDJli9oT88-fTEscyVIdfKa?dl=0)
- [ picture\_as\_pdfPaper](/sites/g/files/omnuum7696/files/CSpaper.pdf)
- [ picture\_as\_pdfAppendix](/sites/g/files/omnuum7696/files/180918appendix.pdf)
 
- [ descriptionPublisher's Version](https://www.dropbox.com/sh/xpfzjx5pzfzgktv/AABDJli9oT88-fTEscyVIdfKa?dl=0)
- [ picture\_as\_pdfPaper](/sites/g/files/omnuum7696/files/CSpaper.pdf)
- [ picture\_as\_pdfAppendix](/sites/g/files/omnuum7696/files/180918appendix.pdf)
 
 

Shen, Zejiang, Kaixuan Zhang, and Melissa Dell. “[A Large Dataset of Historical Japanese Documents With Complex Layouts](/publications/large-dataset-historical-japanese-documents-complex-layouts)”. *IEEE/CVF/Conference/on/Computer/Vision/and/Pattern/Recognition/Workshops*, 2020, 548-59.



 

 

Shen, Zejiang, Kaixuan Zhang, and Melissa Dell. “[A Large Dataset of Historical Japanese Documents With Complex Layouts](/publications/large-dataset-historical-japanese-documents-complex-layouts)”. *IEEE/CVF/Conference/on/Computer/Vision/and/Pattern/Recognition/Workshops*, 2020, 548-59.



 

 

 

- add\_circle do\_not\_disturb\_on Abstract
- [ descriptionPublisher's Version](https://dell-research-harvard.github.io/HJDataset/)
- [ picture\_as\_pdfPaper](/sites/g/files/omnuum7696/files/dell/files/2004.08686.pdf)
 
> Deep learning-based approaches for automatic document layout analysis and content extraction have the potential to unlock rich information trapped in historical documents on a large scale. One major hurdle is the lack of large datasets for training...



 

 

- [ descriptionPublisher's Version](https://dell-research-harvard.github.io/HJDataset/)
- [ picture\_as\_pdfPaper](/sites/g/files/omnuum7696/files/dell/files/2004.08686.pdf)
 
 

 



 

 

 

 - Previous page chevron\_left
- [1](?page=0 "Current page")
- [2](?page=1 "Go to page 2")
- [ Next page chevron\_right ](?page=1 "Go to next page")