Source author record

Brendan O'Connor

Brendan O'Connor appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

16works
12topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

16 published item(s)

preprint2026arXiv

The Advanced X-ray Imaging Satellite (AXIS) Community Science Book

The AXIS Community Science Book represents the collective effort of 592 scientists worldwide to define the transformative science enabled by the Advanced X-ray Imaging Satellite (AXIS), a next-generation X-ray mission selected by NASA's Astrophysics Probe Program for Phase A study. AXIS will advance the legacy of high-angular-resolution X-ray astronomy with ~1.5'' imaging over a wide 24' field of view and an order of magnitude greater collecting area than Chandra in the 0.3-12 keV band. Combining sharp imaging, high throughput, and rapid response capabilities, AXIS will open new windows on virtually every aspect of modern astrophysics, exploring the birth and growth of supermassive black holes, the feedback processes that shape galaxies, the life cycles of stars and exoplanet environments, and the nature of compact stellar remnants, supernova remnants, and explosive transients. This book compiles 138 community-contributed science cases developed by five Science Working Groups focused on AGN and supermassive black holes, galaxy evolution and feedback, compact objects and supernova remnants, stellar physics and exoplanets, and time-domain and multi-messenger astrophysics. Together, these studies establish the scientific foundation for next-generation X-ray exploration in the 2030s and highlight strong synergies with facilities of the 2030s, such as JWST, Roman, Rubin/LSST, SKA, ALMA, ngVLA, and next-generation gravitational-wave and neutrino networks.

preprint2023arXiv

Examining Political Rhetoric with Epistemic Stance Detection

Participants in political discourse employ rhetorical strategies -- such as hedging, attributions, or denials -- to display varying degrees of belief commitments to claims proposed by themselves or others. Traditionally, political scientists have studied these epistemic phenomena through labor-intensive manual content analysis. We propose to help automate such work through epistemic stance prediction, drawn from research in computational semantics, to distinguish at the clausal level what is asserted, denied, or only ambivalently suggested by the author or other mentioned entities (belief holders). We first develop a simple RoBERTa-based model for multi-source stance predictions that outperforms more complex state-of-the-art modeling. Then we demonstrate its novel application to political science by conducting a large-scale analysis of the Mass Market Manifestos corpus of U.S. political opinion books, where we characterize trends in cited belief holders -- respected allies and opposed bogeymen -- across U.S. political ideologies.

preprint2022arXiv

ClioQuery: Interactive Query-Oriented Text Analytics for Comprehensive Investigation of Historical News Archives

Historians and archivists often find and analyze the occurrences of query words in newspaper archives, to help answer fundamental questions about society. But much work in text analytics focuses on helping people investigate other textual units, such as events, clusters, ranked documents, entity relationships, or thematic hierarchies. Informed by a study into the needs of historians and archivists, we thus propose ClioQuery, a text analytics system uniquely organized around the analysis of query words in context. ClioQuery applies text simplification techniques from natural language processing to help historians quickly and comprehensively gather and analyze all occurrences of a query word across an archive. It also pairs these new NLP methods with more traditional features like linked views and in-text highlighting to help engender trust in summarization techniques. We evaluate ClioQuery with two separate user studies, in which historians explain how ClioQuery's novel text simplification features can help facilitate historical research. We also evaluate with a separate quantitative comparison study, which shows that ClioQuery helps crowdworkers find and remember historical information. Such results suggest possible new directions for text analytics in other query-oriented settings.

preprint2021arXiv

Kilonova Detectability with Wide-Field Instruments

Kilonovae are ultraviolet, optical, and infrared transients powered by the radioactive decay of heavy elements following a neutron star merger. Joint observations of kilonovae and gravitational waves can offer key constraints on the source of Galactic $r$-process enrichment, among other astrophysical topics. However, robust constraints on heavy element production requires rapid kilonova detection (within $\sim 1$ day of merger) as well as multi-wavelength observations across multiple epochs. In this study, we quantify the ability of 13 wide field-of-view instruments to detect kilonovae, leveraging a large grid of over 900 radiative transfer simulations with 54 viewing angles per simulation. We consider both current and upcoming instruments, collectively spanning the full kilonova spectrum. The Roman Space Telescope has the highest redshift reach of any instrument in the study, observing kilonovae out to $z \sim 1$ within the first day post-merger. We demonstrate that BlackGEM, DECam, GOTO, the Vera C. Rubin Observatory's LSST, ULTRASAT, and VISTA can observe some kilonovae out to $z \sim 0.1$ ($\sim$475 Mpc), while DDOTI, MeerLICHT, PRIME, $Swift$/UVOT, and ZTF are confined to more nearby observations. Furthermore, we provide a framework to infer kilonova ejecta properties following non-detections and explore variation in detectability with these ejecta parameters.

preprint2020arXiv

Constraints on the circumburst environments of short gamma-ray bursts

Observational follow-up of well localized short gamma-ray bursts (SGRBs) has left $20-30\%$ of the population without a coincident host galaxy association to deep optical and NIR limits ($\gtrsim 26$ mag). These SGRBs have been classified as observationally hostless due to their lack of strong host associations. It has been argued that these hostless SGRBs could be an indication of the large distances traversed by the binary neutron star system (due to natal kicks) between its formation and its merger (leading to a SGRB). The distances of GRBs from their host galaxies can be indirectly probed by the surrounding circumburst densities. We show that a lower limit on those densities can be obtained from early afterglow lightcurves. We find that $\lesssim16\%$ of short GRBs in our sample took place at densities $\lesssim10^{-4}$ cm$^{-3}$. These densities represent the expected range of values at distances greater than the host galaxy's virial radii. We find that out of the five SGRBs in our sample that have been found to be observationally hostless, none are consistent with having occurred beyond the virial radius of their birth galaxies. This implies one of two scenarios. Either these observationally hostless SGRBs occurred outside of the half-light radius of their host galaxy, but well within the galactic halo, or in host galaxies at moderate to high redshifts ($z\gtrsim 2$) that were missed by follow-up observations.

preprint2020arXiv

Text and Causal Inference: A Review of Using Text to Remove Confounding from Causal Estimates

Many applications of computational social science aim to infer causal conclusions from non-experimental data. Such observational data often contains confounders, variables that influence both potential causes and potential effects. Unmeasured or latent confounders can bias causal estimates, and this has motivated interest in measuring potential confounders from observed text. For example, an individual's entire history of social media posts or the content of a news article could provide a rich measurement of multiple confounders. Yet, methods and applications for this problem are scattered across different communities and evaluation practices are inconsistent. This review is the first to gather and categorize these examples and provide a guide to data-processing and evaluation decisions. Despite increased attention on adjusting for confounding using text, there are still many open problems, which we highlight in this paper.

preprint2016arXiv

Demographic Dialectal Variation in Social Media: A Case Study of African-American English

Though dialectal language is increasingly abundant on social media, few resources exist for developing NLP tools to handle such language. We conduct a case study of dialectal language in online conversational text by investigating African-American English (AAE) on Twitter. We propose a distantly supervised model to identify AAE-like language from demographics associated with geo-located messages, and we verify that this language follows well-known AAE linguistic phenomena. In addition, we analyze the quality of existing language identification and dependency parsing tools on AAE-like text, demonstrating that they perform poorly on such text compared to text associated with white speakers. We also provide an ensemble classifier for language identification which eliminates this disparity and release a new corpus of tweets containing AAE-like language.

preprint2016arXiv

Visualizing textual models with in-text and word-as-pixel highlighting

We explore two techniques which use color to make sense of statistical text models. One method uses in-text annotations to illustrate a model's view of particular tokens in particular documents. Another uses a high-level, "words-as-pixels" graphic to display an entire corpus. Together, these methods offer both zoomed-in and zoomed-out perspectives into a model's understanding of text. We show how these interconnected methods help diagnose a classifier's poor performance on Twitter slang, and make sense of a topic model on historical political texts.

preprint2015arXiv

Posterior calibration and exploratory analysis for natural language processing models

Many models in natural language processing define probabilistic distributions over linguistic structures. We argue that (1) the quality of a model' s posterior distribution can and should be directly evaluated, as to whether probabilities correspond to empirical frequencies, and (2) NLP uncertainty can be projected not only to pipeline components, but also to exploratory data analysis, telling a user when to trust and not trust the NLP analysis. We present a method to analyze calibration, and apply it to compare the miscalibration of several commonly used models. We also contribute a coreference sampling algorithm that can create confidence intervals for a political event extraction task.

preprint2014arXiv

Diffusion of Lexical Change in Social Media

Computer-mediated communication is driving fundamental changes in the nature of written language. We investigate these changes by statistical analysis of a dataset comprising 107 million Twitter messages (authored by 2.7 million unique user accounts). Using a latent vector autoregressive model to aggregate across thousands of words, we identify high-level patterns in diffusion of linguistic change over the United States. Our model is robust to unpredictable changes in Twitter's sampling rate, and provides a probabilistic characterization of the relationship of macro-scale linguistic influence to a set of demographic and geographic predictors. The results of this analysis offer support for prior arguments that focus on geographical proximity and population size. However, demographic similarity -- especially with regard to race -- plays an even more central role, as cities with similar racial demographics are far more likely to share linguistic influence. Rather than moving towards a single unified "netspeak" dialect, language evolution in computer-mediated communication reproduces existing fault lines in spoken American English.

preprint2013arXiv

"Groupware for Groups": Problem-Driven Design in Deme

Design choices can be clarified when group interaction software is directed at solving the interaction needs of particular groups that pre-date the groupware. We describe an example: the Deme platform for online deliberation. Traditional threaded conversation systems are insufficient for solving the problem at which Deme is aimed, namely, that the democratic process in grassroots community groups is undermined both by the limited availability of group members for face-to-face meetings and by constraints on the use of information in real-time interactions. We describe and motivate design elements, either implemented or planned for Deme, that addresses this problem. We believe that "problem focused" design of software for preexisting groups provides a useful framework for evaluating the appropriateness of design elements in groupware generally.

preprint2013arXiv

A framework for (under)specifying dependency syntax without overloading annotators

We introduce a framework for lightweight dependency syntax annotation. Our formalism builds upon the typical representation for unlabeled dependencies, permitting a simple notation and annotation workflow. Moreover, the formalism encourages annotators to underspecify parts of the syntax if doing so would streamline the annotation process. We demonstrate the efficacy of this annotation on three languages and develop algorithms to evaluate and compare underspecified annotations.

preprint2013arXiv

An Online Environment for Democratic Deliberation: Motivations, Principles, and Design

We have created a platform for online deliberation called Deme (which rhymes with 'team'). Deme is designed to allow groups of people to engage in collaborative drafting, focused discussion, and decision making using the Internet. The Deme project has evolved greatly from its beginning in 2003. This chapter outlines the thinking behind Deme's initial design: our motivations for creating it, the principles that guided its construction, and its most important design features. The version of Deme described here was written in PHP and was deployed in 2004 and used by several groups (including organizers of the 2005 Online Deliberation Conference). Other papers describe later developments in the Deme project (see Davies et al. 2005, 2008; Davies and Mintz 2009).

preprint2013arXiv

ARKref: a rule-based coreference resolution system

ARKref is a tool for noun phrase coreference. It is a deterministic, rule-based system that uses syntactic information from a constituent parser, and semantic information from an entity recognition component. Its architecture is based on the work of Haghighi and Klein (2009). ARKref was originally written in 2009. At the time of writing, the last released version was in March 2011. This document describes that version, which is open-source and publicly available at: http://www.ark.cs.cmu.edu/ARKref

preprint2013arXiv

Displaying Asynchronous Reactions to a Document: Two Goals and a Design

We describe and motivate three goals for the screen display of asynchronous text deliberation pertaining to a document: (1) visibility of relationships between comments and the text they reference, between different comments, and between group members and the document and discussion, and (2) distinguishability of boundaries between contextually related and unrelated text and comments and between individual authors of documents and comments. Interfaces for document-centered discussion generally fail to fulfill one or both of these goals as well as they could. We describe the design of the new version of Deme, a Web-based platform for online deliberation, and argue that it achieves the two goals better than other recent designs.

preprint2013arXiv

Learning Frames from Text with an Unsupervised Latent Variable Model

We develop a probabilistic latent-variable model to discover semantic frames---types of events and their participants---from corpora. We present a Dirichlet-multinomial model in which frames are latent categories that explain the linking of verb-subject-object triples, given document-level sparsity. We analyze what the model learns, and compare it to FrameNet, noting it learns some novel and interesting frames. This document also contains a discussion of inference issues, including concentration parameter learning; and a small-scale error analysis of syntactic parsing accuracy.