Large Scale Genealogical Information Extraction From Handwritten Quebec Parish Records

Tarride, Solène; Maarand, Martin; Boillet, Mélodie; McGrath, James; Capel, Eugénie; Vézina, Hélène; Kermorvant, Christopher

doi:10.21203/rs.3.rs-2260181/v1

Cited by 3 publications

(3 citation statements)

References 46 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…The aforementioned datasets and their corresponding projects center around two primary challenges: analysing document layout and recognizing handwritten text. On another hand, while not being a table dataset, SIMARA [33] is a dataset of handwritten archive finding aids, comprised of metadata describing historical archives. Finding aids are handwritten and feature the same scientific challenges regarding handwritten text recognition.…”

Section: Historical Index Table Dataset: Paresmentioning

confidence: 99%

“…We aim at localizing text regions to perform text extraction and named entity recognition before indexing them. This is where the added value resides for archives and digital libraries [33], as it can potentially serve as finding aids for census tables and directly contribute to demographic studies. As a result, the detection of text lines is the baseline task we perform on this dataset.…”

Section: Document Layout and Annotationsmentioning

confidence: 99%

“…The detection of text lines has been widely explored in historical manuscript text books [26,9] and other historical documents of different natures, such as newspapers [25], meteorological tables [1] finding aids [33], as well as many other supports. With index tables, one can consider the issue as a two-class image segmentation task: we separate text lines from the background.…”

Section: Document Image Analysismentioning

confidence: 99%

See 2 more Smart Citations

Text Line Detection in Historical Index Tables: Evaluations on a New French PArish REcord Survey Dataset (PARES)

Bernard,

Wall,

Boillet

et al. 2023

Lecture Notes in Computer Science

View full text Add to dashboard Cite

In this paper, we address the challenge of document image analysis for historical index table documents with handwritten records. Demographic studies can gain insight from the use of automatic document analysis in such documents through the study of population movements. To evaluate the efficacy of automatic layout analysis tools, we release the PARES dataset [6], which contains 250 labeled index table images originating from French archives. Also, we run state-of-the-art algorithms (U-FCN, R-CNN and Transformers) in order to detect the lines within index tables, a common prerequisite for handwritten text recognition (HTR). Our results indicate that text line extraction works well with the U-FCN model, while also indicating that Transformer architectures show promise for accurate text line detection in such historical documents with great efficiency. This is a encouraging step towards a Transformer-based architecture for both layout and content detection. This process and dataset represent a first step to automatically analyze handwritten and historical index tables. In addition to this paper and the PARES [6] dataset of historical index tables of 250 images, we release segmentation masks, the code we used to train and test the models, and the models themselves.

show abstract

Section: Historical Index Table Dataset: Paresmentioning

confidence: 99%

Section: Document Layout and Annotationsmentioning

confidence: 99%

Section: Document Image Analysismentioning

confidence: 99%

See 1 more Smart Citation

Text Line Detection in Historical Index Tables: Evaluations on a New French PArish REcord Survey Dataset (PARES)

Bernard,

Wall,

Boillet

et al. 2023

Lecture Notes in Computer Science

View full text Add to dashboard Cite

show abstract

Key-Value Information Extraction from Full Handwritten Pages

Tarride,

Boillet,

Kermorvant

2023

Lecture Notes in Computer Science

View full text Add to dashboard Cite

Segmenting large historical notarial manuscripts into multi-page deeds

Prieto,

Becerra,

Toselli

et al. 2024

Pattern Anal Applic

View full text Add to dashboard Cite

Archives around the world hold vast digitized series of historical manuscript books or “bundles” containing, among others, notarial records also known as “deeds” or “acts”. One of the first steps to provide metadata which describe the contents of those bundles is to segment them into their individual deeds. Even if deeds are often page-aligned, as in the bundles considered in the present work, this is a time-consuming task, often prohibitive given the huge scale of the manuscript series involved. Unlike traditional Layout Analysis methods for page-level segmentation, our approach goes beyond the realm of a single-page image, providing consistent deed detection results on full bundles. This is achieved in two tightly integrated steps: first, we estimate the class-posterior at the page level for the “initial”, “middle”, and “final” classes; then we “decode” these posteriors applying a series of sequentiality consistency constraints to obtain a consistent book segmentation. Experiments are presented for four large historical manuscripts, varying the number of “deeds” used for training. Two metrics are introduced to assess the quality of book segmentation, one of them taking into account the loss of information entailed by segmentation errors. The problem formalization, the metrics and the empirical work significantly extend our previous works on this topic.

show abstract

Large Scale Genealogical Information Extraction From Handwritten Quebec Parish Records

Cited by 3 publications

References 46 publications

Text Line Detection in Historical Index Tables: Evaluations on a New French PArish REcord Survey Dataset (PARES)

Text Line Detection in Historical Index Tables: Evaluations on a New French PArish REcord Survey Dataset (PARES)

Key-Value Information Extraction from Full Handwritten Pages

Segmenting large historical notarial manuscripts into multi-page deeds

Contact Info

Product

Resources

About