Source author record

Li Jin

Li Jin appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

23works
15topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

23 published item(s)

preprint2026arXiv

ES-Mem: Event Segmentation-Based Memory for Long-Term Dialogue Agents

Memory is critical for dialogue agents to maintain coherence and enable continuous adaptation in long-term interactions. While existing memory mechanisms offer basic storage and retrieval capabilities, they are hindered by two primary limitations: (1) rigid memory granularity often disrupts semantic integrity, resulting in fragmented and incoherent memory units; (2) prevalent flat retrieval paradigms rely solely on surface-level semantic similarity, neglecting the structural cues of discourse required to navigate and locate specific episodic contexts. To mitigate these limitations, drawing inspiration from Event Segmentation Theory, we propose ES-Mem, a framework incorporating two core components: (1) a dynamic event segmentation module that partitions long-term interactions into semantically coherent events with distinct boundaries; (2) a hierarchical memory architecture that constructs multi-layered memories and leverages boundary semantics to anchor specific episodic memory for precise context localization. Evaluations on two memory benchmarks demonstrate that ES-Mem yields consistent performance gains over baseline methods. Furthermore, the proposed event segmentation module exhibits robust applicability on dialogue segmentation datasets.

preprint2023arXiv

Evaluation of Public Transit Systems under Short Random Service Suspensions: A Bulk-Service Queuing Approach

This paper proposes a stochastic framework to evaluate the performance of public transit systems under short random service suspensions. We aim to derive closed-form formulations of the mean and variance of the queue length and waiting time. A bulk-service queue model is adopted to formulate the queuing behavior in the system. The random service suspension is modeled as a two-state (disruption and normal) Markov process. We prove that headway is distributed as the difference between two compound Poisson exponential random variables. The distribution is used to specify the mean and variance of queue length and waiting time at each station with analytical formulations. The closed-form stability condition of the system is also derived, implying that the system is more likely to be unstable with high incident rates and long incident duration. The proposed model is implemented on a bus network. Results show that higher incident rates and higher average incident duration will increase both the mean and variance of queue length and waiting time, which are consistent with the theoretical analysis. Crowding stations are more vulnerable to random service suspensions. The theoretical results are validated with a simulation model, showing consistency between the two outcomes.

preprint2022arXiv

Resilient Ramp Control for Highways Facing Stochastic Perturbations

Highway capacity is often subject to stochastic perturbations due to the combined effects of weather, traffic mixture, driver behavior, etc. This paper is motivated by the need of a systematic approach to traffic control with performance guarantees in the face of such perturbations. We develop a novel control-theoretic method for designing perturbation-resilient ramp metering. We consider a cell-transmission model with 1) Markovian cell capacities and 2) buffers representing on-ramps and upstream mainline. Using this model, we analyze the stability of on-ramp queues by constructing piecewise Lyapunov functions that consider the nature of nonlinear traffic dynamics. Then, we design ramp controllers that guarantee bounds for throughput and queue sizes. We also formulate the problem of coordinated ramp metering as a bi-level optimization with non-convex inner sub-problems. To address the computational issue in solving this problem, we also consider localized and partially coordinated reformulations. A case study of a 18.1-km highway in Los Angeles, USA indicates a 8.3\% (resp. 9.9\%) reduction of vehicle-hours-traveled obtained by the localized (resp. partially coordinated) control, both outperforming the classical ALINEA and METALINE controllers.

preprint2021arXiv

Dynamic Queue-Jump Lane for Emergency Vehicles under Partially Connected Settings: A Multi-Agent Deep Reinforcement Learning Approach

Emergency vehicle (EMV) service is a key function of cities and is exceedingly challenging due to urban traffic congestion. The main reason behind EMV service delay is the lack of communication and cooperation between vehicles blocking EMVs. In this paper, we study the improvement of EMV service under V2X connectivity. We consider the establishment of dynamic queue jump lanes (DQJLs) based on real-time coordination of connected vehicles in the presence of non-connected human-driven vehicles. We develop a novel Markov decision process formulation for the DQJL coordination strategies, which explicitly accounts for the uncertainty of drivers' yielding pattern to approaching EMVs. Based on pairs of neural networks representing actors and critics for agent vehicles, we develop a multi-agent actor-critic deep reinforcement learning algorithm that handles a varying number of vehicles and a random proportion of connected vehicles in the traffic. Approaching the optimal coordination strategies via indirect and direct reinforcement learning, we present two schemata to address multi-agent reinforcement learning on this connected vehicle application. Both approaches are validated, on a micro-simulation testbed SUMO, to establish a DQJL fast and safely. Validation results reveal that, with DQJL coordination strategies, it saves up to 30% time for EMVs to pass a link-level intelligent urban roadway than the baseline scenario.

preprint2020arXiv

Artificial Intelligence Forecasting of Covid-19 in China

BACKGROUND An alternative to epidemiological models for transmission dynamics of Covid-19 in China, we propose the artificial intelligence (AI)-inspired methods for real-time forecasting of Covid-19 to estimate the size, lengths and ending time of Covid-19 across China. METHODS We developed a modified stacked auto-encoder for modeling the transmission dynamics of the epidemics. We applied this model to real-time forecasting the confirmed cases of Covid-19 across China. The data were collected from January 11 to February 27, 2020 by WHO. We used the latent variables in the auto-encoder and clustering algorithms to group the provinces/cities for investigating the transmission structure. RESULTS We forecasted curves of cumulative confirmed cases of Covid-19 across China from Jan 20, 2020 to April 20, 2020. Using the multiple-step forecasting, the estimated average errors of 6-step, 7-step, 8-step, 9-step and 10-step forecasting were 1.64%, 2.27%, 2.14%, 2.08%, 0.73%, respectively. We predicted that the time points of the provinces/cities entering the plateau of the forecasted transmission dynamic curves varied, ranging from Jan 21 to April 19, 2020. The 34 provinces/cities were grouped into 9 clusters. CONCLUSIONS The accuracy of the AI-based methods for forecasting the trajectory of Covid-19 was high. We predicted that the epidemics of Covid-19 will be over by the middle of April. If the data are reliable and there are no second transmissions, we can accurately forecast the transmission dynamics of the Covid-19 across the provinces/cities in China. The AI-inspired methods are a powerful tool for helping public health planning and policymaking.

preprint2020arXiv

Coordinating Vehicle Platoons for Highway Bottleneck Decongestion and Throughput Improvement

Truck platooning is a technology that is expected to become widespread in the coming years. Apart from the numerous benefits that it brings, its potential effects on the overall traffic situation need to be studied further, especially at bottlenecks and ramps. Assuming we can control the platoons from the infrastructure, they can be used as controlled moving bottlenecks, actuating control actions on the rest of the traffic, and potentially improving the throughput of the whole system. In this paper, we use a multi-class cell transmission model to capture the interaction between truck platoons and background traffic, and propose a corresponding queuing model, which we use for control design. We use platoon speeds, and the number of lanes platoons occupy as control inputs, and design a control strategy for throughput improvement of a highway section with a bottleneck. By postponing and shaping the inflow to the bottleneck, we are able to avoid traffic breakdown and capacity drop, which significantly reduces the total time spent of all vehicles. We derived the estimated improvement in throughput that is achieved by applying the proposed control law, and then tested it in a simulation study and found that the median delay of all vehicles by 75.6% compared to the uncontrolled case. Notably, although they are slowed down while actuating control actions, platooned vehicles experience less delay compared to the case without control, since they avoid going through congestion at the bottleneck.

preprint2020arXiv

Forecasting and evaluating intervention of Covid-19 in the World

When the Covid-19 pandemic enters dangerous new phase, whether and when to take aggressive public health interventions to slow down the spread of COVID-19. To develop the artificial intelligence (AI) inspired methods for real-time forecasting and evaluating intervention strategies to curb the spread of Covid-19 in the World. A modified auto-encoder for modeling the transmission dynamics of the epidemics is developed and applied to the surveillance data of cumulative and new Covid-19 cases and deaths from WHO, as of March 16, 2020. The average errors of 5-step forecasting were 2.5%. The total peak number of cumulative cases and new cases, and the maximum number of cumulative cases in the world with later intervention (comprehensive public health intervention is implemented 4 weeks later) could reach 75,249,909, 10,086,085, and 255,392,154, respectively. The case ending time was January 10, 2021. However, the total peak number of cumulative cases and new cases and the maximum number of cumulative cases in the world with one week later intervention were reduced to 951,799, 108,853 and 1,530,276, respectively. Duration time of the Covid-19 spread would be reduced from 356 days to 232 days. The case ending time was September 8, 2020. We observed that delaying intervention for one month caused the maximum number of cumulative cases to increase 166.89 times, and the number of deaths increase from 53,560 to 8,938,725. We will face disastrous consequences if immediate action to intervene is not taken.

preprint2020arXiv

Resilience of Dynamic Routing in the Face of Recurrent and Random Sensing Faults

Feedback dynamic routing is a commonly used control strategy in transportation systems. This class of control strategies relies on real-time information about the traffic state in each link. However, such information may not always be observable due to temporary sensing faults. In this article, we consider dynamic routing over two parallel routes, where the sensing on each link is subject to recurrent and random faults. The faults occur and clear according to a finite-state Markov chain. When the sensing is faulty on a link, the traffic state on that link appears to be zero to the controller. Building on the theories of Markov processes and monotone dynamical systems, we derive lower and upper bounds for the resilience score, i.e. the guaranteed throughput of the network, in the face of sensing faults by establishing stability conditions for the network. We use these results to study how a variety of key parameters affect the resilience score of the network. The main conclusions are: (i) Sensing faults can reduce throughput and destabilize a nominally stable network; (ii) A higher failure rate does not necessarily reduce throughput, and there may exist a worst rate that minimizes throughput; (iii) Higher correlation between the failure probabilities of two links leads to greater throughput; (iv) A large difference in capacity between two links can result in a drop in throughput.

preprint2020arXiv

SRQA: Synthetic Reader for Factoid Question Answering

The question answering system can answer questions from various fields and forms with deep neural networks, but it still lacks effective ways when facing multiple evidences. We introduce a new model called SRQA, which means Synthetic Reader for Factoid Question Answering. This model enhances the question answering system in the multi-document scenario from three aspects: model structure, optimization goal, and training method, corresponding to Multilayer Attention (MA), Cross Evidence (CE), and Adversarial Training (AT) respectively. First, we propose a multilayer attention network to obtain a better representation of the evidences. The multilayer attention mechanism conducts interaction between the question and the passage within each layer, making the token representation of evidences in each layer takes the requirement of the question into account. Second, we design a cross evidence strategy to choose the answer span within more evidences. We improve the optimization goal, considering all the answers' locations in multiple evidences as training targets, which leads the model to reason among multiple evidences. Third, adversarial training is employed to high-level variables besides the word embedding in our model. A new normalization method is also proposed for adversarial perturbations so that we can jointly add perturbations to several target variables. As an effective regularization method, adversarial training enhances the model's ability to process noisy data. Combining these three strategies, we enhance the contextual representation and locating ability of our model, which could synthetically extract the answer span from several evidences. We perform SRQA on the WebQA dataset, and experiments show that our model outperforms the state-of-the-art models (the best fuzzy score of our model is up to 78.56%, with an improvement of about 2%).

preprint2015arXiv

A New Statistical Framework for Genetic Pleiotropic Analysis of High Dimensional Phenotype Data

The widely used genetic pleiotropic analysis of multiple phenotypes are often designed for examining the relationship between common variants and a few phenotypes. They are not suited for both high dimensional phenotypes and high dimensional genotype (next-generation sequencing) data. To overcome these limitations, we develop sparse structural equation models (SEMs) as a general framework for a new paradigm of genetic analysis of multiple phenotypes. To incorporate both common and rare variants into the analysis, we extend the traditional multivariate SEMs to sparse functional SEMs. To deal with high dimensional phenotype and genotype data, we employ functional data analysis and the alternative direction methods of multiplier (ADMM) techniques to reduce data dimension and improve computational efficiency. Using large scale simulations we showed that the proposed methods have higher power to detect true causal genetic pleiotropic structure than other existing methods. Simulations also demonstrate that the gene-based pleiotropic analysis has higher power than the single variant-based pleiotropic analysis. The proposed method is applied to exome sequence data from the NHLBI Exome Sequencing Project (ESP) with 11 phenotypes, which identifies a network with 137 genes connected to 11 phenotypes and 341 edges. Among them, 114 genes showed pleiotropic genetic effects and 45 genes were reported to be associated with phenotypes in the analysis or other cardiovascular disease (CVD) related phenotypes in the literature.

preprint2015arXiv

Genetic structure of Sino-Tibetan populations revealed by forensic STR loci

The origin and diversification of Sino-Tibetan populations have been a long-standing hot debate. However, the limited genetic information of Tibetan populations keeps this topic far from clear. In the present study, we genotyped 15 forensic autosomal STRs from 803 unrelated Tibetan individuals from Gansu Province (635 from Gannan and 168 from Tianzhu). We combined these data with published dataset to infer a detailed population affinities and admixture of Sino-Tibetan populations. Our results revealed that the genetic structure of Sino-Tibetan populations was strongly correlated with linguistic affiliations. Although the among-population variances are relatively small, the genetic components for Tibetan, Lolo-Burmese, and Han Chinese were quite distinctive, especially for the Deng, Nu, and Derung of Lolo-Burmese. Southern indigenous populations, such as Tai-Kadai and Hmong-Mien populations might have made substantial genetic contribution to Han Chinese and Altaic populations, but not to Tibetans. Likewise, Han Chinese but not Tibetan shared very similar genetic makeups with Altaic populations, which did not support the North Asian origin of Tibetan populations. The dataset generated here are also valuable for forensic identification.

preprint2015arXiv

Random Bits Regression: a Strong General Predictor for Big Data

To improve accuracy and speed of regressions and classifications, we present a data-based prediction method, Random Bits Regression (RBR). This method first generates a large number of random binary intermediate/derived features based on the original input matrix, and then performs regularized linear/logistic regression on those intermediate/derived features to predict the outcome. Benchmark analyses on a simulated dataset, UCI machine learning repository datasets and a GWAS dataset showed that RBR outperforms other popular methods in accuracy and robustness. RBR (available on https://sourceforge.net/projects/rbr/) is very fast and requires reasonable memories, therefore, provides a strong, robust and fast predictor in the big data era.

preprint2015arXiv

The dichotomy structure of Y chromosome Haplogroup N

Haplogroup N-M231 of human Y chromosome is a common clade from Eastern Asia to Northern Europe, being one of the most frequent haplogroups in Altaic and Uralic-speaking populations. Using newly discovered bi-allelic markers from high-throughput DNA sequencing, we largely improved the phylogeny of Haplogroup N, in which 16 subclades could be identified by 33 SNPs. More than 400 males belonging to Haplogroup N in 34 populations in China were successfully genotyped, and populations in Northern Asia and Eastern Europe were also compared together. We found that all the N samples were typed as inside either clade N1-F1206 (including former N1a-M128, N1b-P43 and N1c-M46 clades), most of which were found in Altaic, Uralic, Russian and Chinese-speaking populations, or N2-F2930, common in Tibeto-Burman and Chinese-speaking populations. Our detailed results suggest that Haplogroup N developed in the region of China since the final stage of late Paleolithic Era.

preprint2014arXiv

A brain-wide association study of DISC1 genetic variants reveals a relationship with the structure and functional connectivity of the precuneus in schizophrenia

The Disrupted in Schizophrenia Gene 1 (DISC1) plays a role in both neural signalling and development and is associated with schizophrenia, although its links to altered brain structure and function in this disorder are not fully established. Here we have used structural and functional MRI to investigate links with six DISC1 single nucleotide polymorphisms (SNPs). We employed a brain-wide association analysis (BWAS) together with a Jacknife internal validation approach in 46 schizophrenia patients and 24 matched healthy control subjects. Results from structural MRI showed significant associations between all six DISC1 variants and gray matter volume in the precuneus, post-central gyrus and middle cingulate gyrus. Associations with specific SNPs were found for rs2738880 in the left precuneus and right post-central gyrus, and rs1535530 in the right precuneus and middle cingulate gyrus. Using regions showing structural associations as seeds a resting-state functional connectivity analysis revealed significant associations between all 6 SNPS and connectivity between the right precuneus and inferior frontal gyrus. The connection between the right precuneus and inferior frontal gyrus was also specifically associated with rs821617. Importantly schizophrenia patients showed positive correlations between the six DISC-1 SNPs associated gray matter volume in the left precuneus and right post-central gyrus and negative symptom severity. No correlations with illness duration were found. Our results provide the first evidence suggesting a key role for structural and functional connectivity associations between DISC1 polymorphisms and the precuneus in schizophrenia.

preprint2014arXiv

Genome-wide Scan of Archaic Hominin Introgressions in Eurasians Reveals Complex Admixture History

Introgressions from Neanderthals and Denisovans were detected in modern humans. Introgressions from other archaic hominins were also implicated, however, identification of which poses a great technical challenge. Here, we introduced an approach in identifying introgressions from all possible archaic hominins in Eurasian genomes, without referring to archaic hominin sequences. We focused on mutations emerged in archaic hominins after their divergence from modern humans (denoted as archaic-specific mutations), and identified introgressive segments which showed significant enrichment of archaic-specific mutations over the rest of the genome. Furthermore, boundaries of introgressions were identified using a dynamic programming approach to partition whole genome into segments which contained different levels of archaic-specific mutations. We found that detected introgressions shared more archaic-specific mutations with Altai Neanderthal than they shared with Denisovan, and 60.3% of archaic hominin introgressions were from Neanderthals. Furthermore, we detected more introgressions from two unknown archaic hominins whom diverged with modern humans approximately 859 and 3,464 thousand years ago. The latter unknown archaic hominin contributed to the genomes of the common ancestors of modern humans and Neanderthals. In total, archaic hominin introgressions comprised 2.4% of Eurasian genomes. Above results suggested a complex admixture history among hominins. The proposed approach could also facilitate admixture research across species.

preprint2013arXiv

Agriculture driving male expansion in Neolithic Time

The emergence of agriculture is suggested to have driven extensive human population growths. However, genetic evidence from maternal mitochondrial genomes suggests major population expansions began before the emergence of agriculture. Therefore, role of agriculture that played in initial population expansions still remains controversial. Here, we analyzed a set of globally distributed whole Y chromosome and mitochondrial genomes of 526 male samples from 1000 Genome Project. We found that most major paternal lineage expansions coalesced in Neolithic Time. The estimated effective population sizes through time revealed strong evidence for 10- to 100-fold increase in population growth of males with the advent of agriculture. This sex-biased Neolithic expansion might result from the reduction in hunting-related mortality of males.

preprint2013arXiv

Convergence of Y chromosome STR haplotypes from different SNP haplogroups compromises accuracy of haplogroup prediction

Short tandem repeats (STRs) and single nucleotide polymorphisms (SNPs) are two kinds of commonly used markers in Y chromosome studies of forensic and population genetics. There has been increasing interest in the cost saving strategy by using the STR haplotypes to predict SNP haplogroups. However, the convergence of Y chromosome STR haplotypes from different haplogroups might compromise the accuracy of haplogroup prediction. Here, we compared the worldwide Y chromosome lineages at both haplogroup level and haplotype level to search for the possible haplotype similarities among haplogroups. The similar haplotypes between haplogroups B and I2, C1 and E1b1b1, C2 and E1b1a1, H1 and J, L and O3a2c1, O1a and N, O3a1c and O3a2b, and M1 and O3a2 have been found, and those similarities reduce the accuracy of prediction.

preprint2013arXiv

Global patterns of sex-biased migrations in humans

A series of studies have revealed the among-population components of genetic variation are higher for the paternal Y chromosome than for the maternal mitochondrial DNA (mtDNA), which indicates sex-biased migrations in human populations. However, this phenomenon might be also an ascertainment bias due to nonrandom sampling of SNPs. To eliminate the possible bias, we used the whole Y chromosome and mtDNA sequence data of 491 individuals from the 1000 Genomes Project Phase I to address the sex-biased migration dispute. We found that genetic differentiation between populations was higher for Y chromosome than for the mtDNA at global scales. The migration rate of female might be three times higher than that of male, assuming the effective population size is the same for male and female.

preprint2013arXiv

Natural selection on human Y chromosomes

The paternally inherited Y chromosome has been widely used in population genetic studies to understand relationships among human populations. Our interpretation of Y chromosomal evidence about population history and genetics has rested on the assumption that all the Y chromosomal markers in the male-specific region (MSY) are selectively neutral. However, the very low diversity of Y chromosome has drawn a long debate about whether natural selection has affected this chromosome or not. In recent several years, the progress in Y chromosome sequencing has helped to address this dispute. Purifying selection has been detected in the X-degenerate genes of human Y chromosomes and positive selection might also have an influence in the evolution of testis-related genes in the ampliconic regions. Those new findings remind us to take the effect of natural selection into account when we use Y chromosome in population genetic studies.

preprint2013arXiv

Present Y chromosomes support the Persian ancestry of Sayyid Ajjal Shams al-Din Omar and Eminent Navigator Zheng He

Sayyid Ajjal is the ancestor of many Muslims in areas all across China. And one of his descendants is the famous Navigator of Ming Dynasty, Zheng He, who led the largest armada in the world of 15th century. The origin of Sayyid Ajjal's family remains unclear although many studies have been done on this topic of Muslim history. In this paper, we studied the Y chromosomes of his present descendants, and found they all have haplogroup L1a-M76, proving a southern Persian origin.

preprint2013arXiv

Y Chromosomes of 40% Chinese Are Descendants of Three Neolithic Super-grandfathers

Demographic change of human populations is one of the central questions for delving into the past of human beings. To identify major population expansions related to male lineages, we sequenced 78 East Asian Y chromosomes at 3.9 Mbp of the non-recombining region (NRY), discovered >4,000 new SNPs, and identified many new clades. The relative divergence dates can be estimated much more precisely using molecular clock. We found that all the Paleolithic divergences were binary; however, three strong star-like Neolithic expansions at ~6 kya (thousand years ago) (assuming a constant substitution rate of 1e-9/bp/year) indicates that ~40% of modern Chinese are patrilineal descendants of only three super-grandfathers at that time. This observation suggests that the main patrilineal expansion in China occurred in the Neolithic Era and might be related to the development of agriculture.

preprint2012arXiv

The GenoChip: A New Tool for Genetic Anthropology

The Genographic Project is an international effort using genetic data to chart human migratory history. The project is non-profit and non-medical, and through its Legacy Fund supports locally led efforts to preserve indigenous and traditional cultures. In its second phase, the project is focusing on markers from across the entire genome to obtain a more complete understanding of human genetic variation. Although many commercial arrays exist for genome-wide SNP genotyping, they were designed for medical genetic studies and contain medically related markers that are not appropriate for global population genetic studies. GenoChip, the Genographic Project's new genotyping array, was designed to resolve these issues and enable higher-resolution research into outstanding questions in genetic anthropology. We developed novel methods to identify AIMs and genomic regions that may be enriched with alleles shared with ancestral hominins. Overall, we collected and ascertained AIMs from over 450 populations. Containing an unprecedented number of Y-chromosomal and mtDNA SNPs and over 130,000 SNPs from the autosomes and X-chromosome, the chip was carefully vetted to avoid inclusion of medically relevant markers. The GenoChip results were successfully validated. To demonstrate its capabilities, we compared the FST distributions of GenoChip SNPs to those of two commercial arrays for three continental populations. While all arrays yielded similarly shaped (inverse J) FST distributions, the GenoChip autosomal and X-chromosomal distributions had the highest mean FST, attesting to its ability to discern subpopulations. The GenoChip is a dedicated genotyping platform for genetic anthropology and promises to be the most powerful tool available for assessing population structure and migration history.

preprint2009arXiv

Separation of piezoelectric grain resonance and domain wall dispersion in PZT ceramics

We report on the experimental investigation of a high-frequency (1MHz - 1.8GHz) dielectric dispersion in unpoled and poled Pb(Zr,Ti)O3 ceramics. Two overlapping loss peaks could be revealed in the dielectric spectrum. The linear dependence between the lower-frequency peak position and average grain size D, which holds for D< 10mkm, indicates that the corresponding polarization mechanism originates from piezoelectric resonances of grains. The intensity of the higher-frequency peak is drastically reduced by poling. It is thus proposed that this loss peak is related to domain-wall contribution to the dielectric dispersion.