Source author record

Bhushan Kotnis

Bhushan Kotnis appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

12works
10topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

12 published item(s)

preprint2022arXiv

A Human-Centric Assessment Framework for AI

With the rise of AI systems in real-world applications comes the need for reliable and trustworthy AI. An essential aspect of this are explainable AI systems. However, there is no agreed standard on how explainable AI systems should be assessed. Inspired by the Turing test, we introduce a human-centric assessment framework where a leading domain expert accepts or rejects the solutions of an AI system and another domain expert. By comparing the acceptance rates of provided solutions, we can assess how the AI system performs compared to the domain expert, and whether the AI system's explanations (if provided) are human-understandable. This setup -- comparable to the Turing test -- can serve as a framework for a wide range of human-centric AI system assessments. We demonstrate this by presenting two instantiations: (1) an assessment that measures the classification accuracy of a system with the option to incorporate label uncertainties; (2) an assessment where the usefulness of provided explanations is determined in a human-centric manner.

preprint2022arXiv

AnnIE: An Annotation Platform for Constructing Complete Open Information Extraction Benchmark

Open Information Extraction (OIE) is the task of extracting facts from sentences in the form of relations and their corresponding arguments in schema-free manner. Intrinsic performance of OIE systems is difficult to measure due to the incompleteness of existing OIE benchmarks: the ground truth extractions do not group all acceptable surface realizations of the same fact that can be extracted from a sentence. To measure performance of OIE systems more realistically, it is necessary to manually annotate complete facts (i.e., clusters of all acceptable surface realizations of the same fact) from input sentences. We propose AnnIE: an interactive annotation platform that facilitates such challenging annotation tasks and supports creation of complete fact-oriented OIE evaluation benchmarks. AnnIE is modular and flexible in order to support different use case scenarios (i.e., benchmarks covering different types of facts). We use AnnIE to build two complete OIE benchmarks: one with verb-mediated facts and another with facts encompassing named entities. Finally, we evaluate several OIE systems on our complete benchmarks created with AnnIE. Our results suggest that existing incomplete benchmarks are overly lenient, and that OIE systems are not as robust as previously reported. We publicly release AnnIE under non-restrictive license.

preprint2022arXiv

BenchIE: A Framework for Multi-Faceted Fact-Based Open Information Extraction Evaluation

Intrinsic evaluations of OIE systems are carried out either manually -- with human evaluators judging the correctness of extractions -- or automatically, on standardized benchmarks. The latter, while much more cost-effective, is less reliable, primarily because of the incompleteness of the existing OIE benchmarks: the ground truth extractions do not include all acceptable variants of the same fact, leading to unreliable assessment of the models' performance. Moreover, the existing OIE benchmarks are available for English only. In this work, we introduce BenchIE: a benchmark and evaluation framework for comprehensive evaluation of OIE systems for English, Chinese, and German. In contrast to existing OIE benchmarks, BenchIE is fact-based, i.e., it takes into account informational equivalence of extractions: our gold standard consists of fact synsets, clusters in which we exhaustively list all acceptable surface forms of the same fact. Moreover, having in mind common downstream applications for OIE, we make BenchIE multi-faceted; i.e., we create benchmark variants that focus on different facets of OIE evaluation, e.g., compactness or minimality of extractions. We benchmark several state-of-the-art OIE systems using BenchIE and demonstrate that these systems are significantly less effective than indicated by existing OIE benchmarks. We make BenchIE (data and evaluation code) publicly available on https://github.com/gkiril/benchie.

preprint2022arXiv

Human-Centric Research for NLP: Towards a Definition and Guiding Questions

With Human-Centric Research (HCR) we can steer research activities so that the research outcome is beneficial for human stakeholders, such as end users. But what exactly makes research human-centric? We address this question by providing a working definition and define how a research pipeline can be split into different stages in which human-centric components can be added. Additionally, we discuss existing NLP with HCR components and define a series of guiding questions, which can serve as starting points for researchers interested in exploring human-centric research approaches. We hope that this work would inspire researchers to refine the proposed definition and to pose other questions that might be meaningful for achieving HCR.

preprint2022arXiv

milIE: Modular & Iterative Multilingual Open Information Extraction

Open Information Extraction (OpenIE) is the task of extracting (subject, predicate, object) triples from natural language sentences. Current OpenIE systems extract all triple slots independently. In contrast, we explore the hypothesis that it may be beneficial to extract triple slots iteratively: first extract easy slots, followed by the difficult ones by conditioning on the easy slots, and therefore achieve a better overall extraction. Based on this hypothesis, we propose a neural OpenIE system, milIE, that operates in an iterative fashion. Due to the iterative nature, the system is also modular -- it is possible to seamlessly integrate rule based extraction systems with a neural end-to-end system, thereby allowing rule based systems to supply extraction slots which milIE can leverage for extracting the remaining slots. We confirm our hypothesis empirically: milIE outperforms SOTA systems on multiple languages ranging from Chinese to Arabic. Additionally, we are the first to provide an OpenIE test dataset for Arabic and Galician.

preprint2021arXiv

Answering Complex Queries in Knowledge Graphs with Bidirectional Sequence Encoders

Representation learning for knowledge graphs (KGs) has focused on the problem of answering simple link prediction queries. In this work we address the more ambitious challenge of predicting the answers of conjunctive queries with multiple missing entities. We propose Bi-Directional Query Embedding (BIQE), a method that embeds conjunctive queries with models based on bi-directional attention mechanisms. Contrary to prior work, bidirectional self-attention can capture interactions among all the elements of a query graph. We introduce a new dataset for predicting the answer of conjunctive query and conduct experiments that show BIQE significantly outperforming state of the art baselines.

preprint2016arXiv

Cost Effective Campaigning in Social Networks

Campaigners are increasingly using online social networking platforms for promoting products, ideas and information. A popular method of promoting a product or even an idea is incentivizing individuals to evangelize the idea vigorously by providing them with referral rewards in the form of discounts, cash backs, or social recognition. Due to budget constraints on scarce resources such as money and manpower, it may not be possible to provide incentives for the entire population, and hence incentives need to be allocated judiciously to appropriate individuals for ensuring the highest possible outreach size. We aim to do the same by formulating and solving an optimization problem using percolation theory. In particular, we compute the set of individuals that are provided incentives for minimizing the expected cost while ensuring a given outreach size. We also solve the problem of computing the set of individuals to be incentivized for maximizing the outreach size for given cost budget. The optimization problem turns out to be non trivial; it involves quantities that need to be computed by numerically solving a fixed point equation. Our primary contribution is, that for a fairly general cost structure, we show that the optimization problems can be solved by solving a simple linear program. We believe that our approach of using percolation theory to formulate an optimization problem is the first of its kind.

preprint2016arXiv

Evaluating the Usefulness of Paratransgenesis for Malaria Control

Malaria is a serious global health problem which is especially devastating to the developing world. Mosquitoes are the carriers of the parasite responsible for the disease, and hence malaria control programs focus on controlling mosquito populations. This is done primarily through the spraying of insecticides, or through the use of insecticide treated bed nets. However, usage of these insecticides exerts massive selection pressure on mosquitoes, resulting in insecticide resistant mosquito breeds. Hence, developing alternative strategies is crucial for sustainable malaria control. Here we explore the usefulness of paratransgenesis, i.e., introducing genetically engineered bacteria which secrete anti-plasmodium molecules, inside the mosquito midgut. The bacteria enter a mosquito's midgut when it drinks from a sugar bait, i.e., a sugar solution containing the bacterium. We formulate a mathematical model for evaluating the number of such baits required for preventing an outbreak. We study scenarios where vectors and hosts mix homogeneously as well as heterogeneously. We perform a full stability analysis and calculate the basic reproductive number for both the cases. Additionally, for the heterogeneous mixing scenario, we propose a targeted bait distribution strategy. The optimal bait allocation is calculated and is found to be extremely efficient in terms of bait usage. Our analyses suggest that paratransgenesis can prevent an outbreak, and hence it offers a viable and sustainable path to malaria control.

preprint2016arXiv

Game Theoretic Analysis of Tree Based Referrals for Crowd Sensing Social Systems with Passive Rewards

Participatory crowd sensing social systems rely on the participation of large number of individuals. Since humans are strategic by nature, effective incentive mechanisms are needed to encourage participation. A popular mechanism to recruit individuals is through referrals and passive incentives such as geometric incentive mechanisms used by the winning team in the 2009 DARPA Network Challenge and in multi level marketing schemes. The effect of such recruitment schemes on the effort put in by recruited strategic individuals is not clear. This paper attempts to fill this gap. Given a referral tree and the direct and passive reward mechanism, we formulate a network game where agents compete for finishing crowd sensing tasks. We characterize the Nash equilibrium efforts put in by the agents and derive closed form expressions for the same. We discover free riding behavior among nodes who obtain large passive rewards. This work has implications on designing effective recruitment mechanisms for crowd sourced tasks. For example, usage of geometric incentive mechanisms to recruit large number of individuals may not result in proportionate effort because of free riding.

preprint2016arXiv

Incentivized Campaigning in Social Networks

Campaigners, advertisers and activists are increasingly turning to social recommendation mechanisms, provided by social media, for promoting their products, services, brands and even ideas. However, many times, such social network based campaigns perform poorly in practice because the intensity of the recommendations drastically reduces beyond a few hops from the source. A natural strategy for maintaining the intensity is to provide incentives. In this paper, we address the problem of minimizing the cost incurred by the campaigner for incentivizing a fraction of individuals in the social network, while ensuring that the campaign message reaches a given expected fraction of individuals. We also address the dual problem of maximizing the campaign penetration for a resource constrained campaigner. To help us understand and solve the above mentioned problems, we use percolation theory to formally state them as optimization problems. These problems are not amenable to traditional approaches because of a fixed point equation that needs to be solved numerically. However, we use results from reliability theory to establish some key properties of the fixed point, which in turn enables us to solve these problems using algorithms that are linearithmic in maximum node degree. Furthermore, we evaluate the efficacy of the analytical solution by performing simulations on real world networks.

preprint2016arXiv

Percolation on Networks with Antagonistic and Dependent Interactions

Drawing inspiration from real world interacting systems we study a system consisting of two networks that exhibit antagonistic and dependent interactions. By antagonistic and dependent interactions, we mean, that a proportion of functional nodes in a network cause failure of nodes in the other, while failure of nodes in the other results in failure of links in the first. As opposed to interdependent networks, which can exhibit first order phase transitions, we find that the phase transitions in such networks are continuous. Our analysis shows that, compared to an isolated network, the system is more robust against random attacks. Surprisingly, we observe a region in the parameter space where the giant connected components of both networks start oscillating. Furthermore, we find that for Erdos-Renyi and scale free networks the system oscillates only when the dependency and antagonism between the two networks is very high. We believe that this study can further our understanding of real world interacting systems.

preprint2014arXiv

Cost Effective Rumor Containment in Social Networks

The spread of rumors through social media and online social networks can not only disrupt the daily lives of citizens but also result in loss of life and property. A rumor spreads when individuals, who are unable decide the authenticity of the information, mistake the rumor as genuine information and pass it on to their acquaintances. We propose a solution where a set of individuals (based on their degree) in the social network are trained and provided resources to help them distinguish a rumor from genuine information. By formulating an optimization problem we calculate the optimum set of individuals, who must undergo training, and the quality of training that minimizes the expected training cost and ensures an upper bound on the size of the rumor outbreak. Our primary contribution is that although the optimization problem turns out to be non convex, we show that the problem is equivalent to solving a set of linear programs. This result also allows us to solve the problem of minimizing the size of rumor outbreak for a given cost budget. The optimum solution displays an interesting pattern which can be implemented as a heuristic. These results can prove to be very useful for social planners and law enforcement agencies for preventing dangerous rumors and misinformation epidemics.