Source author record

Richard Furuta

Richard Furuta appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

2works
7topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

2 published item(s)

preprint2016arXiv

Font Identification in Historical Documents Using Active Learning

Identifying the type of font (e.g., Roman, Blackletter) used in historical documents can help optical character recognition (OCR) systems produce more accurate text transcriptions. Towards this end, we present an active-learning strategy that can significantly reduce the number of labeled samples needed to train a font classifier. Our approach extracts image-based features that exploit geometric differences between fonts at the word level, and combines them into a bag-of-word representation for each page in a document. We evaluate six sampling strategies based on uncertainty, dissimilarity and diversity criteria, and test them on a database containing over 3,000 historical documents with Blackletter, Roman and Mixed fonts. Our results show that a combination of uncertainty and diversity achieves the highest predictive accuracy (89% of test cases correctly classified) while requiring only a small fraction of the data (17%) to be labeled. We discuss the implications of this result for mass digitization projects of historical documents.

preprint2011arXiv

Distributed Collections of Web Pages in the Wild

As the Distributed Collection Manager's work on building tools to support users maintaining collections of changing web-based resources has progressed, questions about the characteristics of people's collections of web pages have arisen. Simultaneously, work in the areas of social bookmarking, social news, and subscription-based technologies have been taking the existence, usage, and utility of this data for granted with neither investigation into what people are doing with their collections nor how they are trying to maintain them. In order to address these concerns, we performed an online user study of 125 individuals from a variety of online and offline communities, such as the reddit social news user community and the graduate student body in our department. From this study we were able to examine a user's needs for a system to manage their web-based distributed collections, how their current tools affect their ability to maintain their collections, and what the characteristics of their current practices and problems in maintaining their web-based collections were. We also present extensions and improvements being made to the system both in order to adapt DCM for usage in the Ensemble project and to meet the requirements found by our user study.