Source author record

Jihee Kim

Jihee Kim appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

5works
6topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

5 published item(s)

preprint2026arXiv

In-Context Examples Suppress Scientific Knowledge Recall in LLMs

Scientific reasoning rarely stops at what is directly observable; it often requires uncovering hidden structure from data. From estimating reaction constants in chemistry to inferring demand elasticities in economics, this latent structure recovery is what distinguishes scientific reasoning from curve fitting. Large language models (LLMs) can often recall and apply relevant scientific formulas, but we show that this ability is surprisingly easy to suppress. We show that adding in-context examples makes models rely less on pretrained domain knowledge, even when those examples are generated by the very same formula. Rather than reinforcing knowledge-driven derivation, examples shift computation toward empirical pattern fitting. We document this knowledge displacement on 60 latent structure recovery tasks across five scientific domains, 6,000 trials, and four models. This displacement is consistent across domains, but its accuracy consequences depend on how the displaced strategy compares to the one that replaces it: the same shift can lower accuracy, leave it unchanged, or appear to improve it. In all cases, however, the model shifts away from knowledge-driven reasoning. For practitioners deploying LLMs on scientific tasks, the message is cautionary: in-context examples may displace, rather than reinforce, the knowledge they are intended to support.

preprint2022arXiv

Learning Economic Indicators by Aggregating Multi-Level Geospatial Information

High-resolution daytime satellite imagery has become a promising source to study economic activities. These images display detailed terrain over large areas and allow zooming into smaller neighborhoods. Existing methods, however, have utilized images only in a single-level geographical unit. This research presents a deep learning model to predict economic indicators via aggregating traits observed from multiple levels of geographical units. The model first measures hyperlocal economy over small communities via ordinal regression. The next step extracts district-level features by summarizing interconnection among hyperlocal economies. In the final step, the model estimates economic indicators of districts via aggregating the hyperlocal and district information. Our new multi-level learning model substantially outperforms strong baselines in predicting key indicators such as population, purchasing power, and energy consumption. The model is also robust against data shortage; the trained features from one country can generalize to other countries when evaluated with data gathered from Malaysia, the Philippines, Thailand, and Vietnam. We discuss the multi-level model's implications for measuring inequality, which is the essential first step in policy and social science research on inequality and poverty.

preprint2022arXiv

Monolithic Active Pixel Sensors on CMOS technologies

Collider detectors have taken advantage of the resolution and accuracy of silicon detectors for at least four decades. Future colliders will need large areas of silicon sensors for low mass trackers and sampling calorimetry. Monolithic Active Pixel Sensors (MAPS), in which Si diodes and readout circuitry are combined in the same pixels, and can be fabricated in some of standard CMOS processes, are a promising technology for high-granularity and light detectors. In this paper we review 1) the requirements on MAPS for trackers and electromagnetic calorimeters (ECal) at future colliders experiments, 2) the ongoing efforts towards dedicated MAPS for the Electron-Ion Collider (EIC) at BNL, for which the EIC Silicon Consortium was already instantiated, and 3) space-born applications for MeV $γ$-ray experiments with MAPS based trackers (AstroPix).

preprint2013arXiv

A comparison study of CORSIKA and COSMOS simulations for extensive air showers

Cosmic rays with energy exceeding ~ 10^{18} eV are referred to as ultra-high energy cosmic rays (UHECRs). Monte Carlo codes for extensive air shower (EAS) simulate the development of EASs initiated by UHECRs in the Earth's atmosphere. Experiments to detect UHECRs utilize EAS simulations to estimate their energy, arrival direction, and composition. In this paper, we compare EAS simulations with two different codes, CORSIKA and COSMOS, presenting quantities including the longitudinal distribution of particles, depth of shower maximum, kinetic energy distribution of particle at the ground, and energy deposited to the air. We then discuss implications of our results to UHECR experiments.

preprint2011arXiv

Comparison of CORSIKA and COSMOS simulations

Ultra-high-energy cosmic rays (UHECRs) refer to cosmic rays with energy above 10^{18} eV. UHECR experiments utilize simulations of extensive air shower to estimate the properties of UHECRs. The Telescope Array (TA) experiment employs the Monte Carlo codes of CORSIKA and COSMOS to obtain EAS simulations. In this paper, we compare the results of the simulations obtained from CORSIKA and COSMOS and report differences between them in terms of the longitudinal distribution, Xmax-value, calorimetric energy, and energy spectrum at ground.