Source author record

Noel Crespi

Noel Crespi appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

17works
14topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

17 published item(s)

preprint2026arXiv

Information Density as a Quantitative Measure for AI-enabled Virtual Sensing: Feasibility and Limits

Modern IoT and sensor networks generate vast amounts of data, posing significant challenges for storage, transmission, and real-time processing. Traditional approaches, such as compressive sensing and machine learning-based compression, often suffer from computational inefficiencies and irreversible data loss. This paper introduces Information Density as a quantitative metric to support sensor deployment and enable AI-driven virtual sensing. We propose a framework that leverages spatial, temporal and inter-modal correlations among sensor signals to perform sensing tasks even in the absence of physical sensors. Two complementary measures: (i) Phase in Eigen Space and (ii) Mutual Information, are developed to quantify and assess information density, enabling the selection of optimal sensor configurations across both intra-modality and cross-modality scenarios. Validated using real-world data from Madrid's smart city infrastructure, this framework demonstrates the feasibility of replacing physical sensors with virtual ones under bounded error conditions (e.g., achieving $<3.21\%$ mean error with a single sensor). The results highlight the potential for scalable and energy-efficient sensing systems in smart environments.

preprint2026arXiv

Med-StepBench: A Hierarchical Reasoning Framework for Evaluating Hallucinations in Medical Vision-Language Models

Large vision-language models (VLMs) demonstrate strong performance in medical image understanding, but frequently generate clinically plausible yet incorrect statements, raising significant safety concerns. Existing medical hallucination benchmarks primarily focus on 2D imaging with one-shot diagnostic questions, offering limited insight into whether predictions are grounded in correct localization and abnormality identification, allowing critical reasoning errors to remain hidden behind seemingly correct diagnoses. We introduce Med-StepBench, the first large-scale benchmark for step-wise hallucination detection in 3D oncological PET/CT, comprising over 12,000 images and more than 1,000,000 image-statement pairs across volumetric and multi-view 2D data, which decomposes clinical reasoning into four expert-designed diagnostic stages. Using clinician-verified annotations, we perform the first step-level evaluation of general-purpose and medical VLMs, revealing systematic failure modes obscured by aggregate accuracy metrics. Furthermore, we show that current VLMs are highly susceptible to adversarial yet clinically plausible intermediate explanations, which significantly amplify hallucinations despite contradictory visual evidence. Together, our findings highlight fundamental limitations in grounding multi-step clinical reasoning and establish Med-StepBench as a rigorous benchmark for developing safer and more reliable medical VLMs.

preprint2022arXiv

BERT-based Ensemble Approaches for Hate Speech Detection

With the freedom of communication provided in online social media, hate speech has increasingly generated. This leads to cyber conflicts affecting social life at the individual and national levels. As a result, hateful content classification is becoming increasingly demanded for filtering hate content before being sent to the social networks. This paper focuses on classifying hate speech in social media using multiple deep models that are implemented by integrating recent transformer-based language models such as BERT, and neural networks. To improve the classification performances, we evaluated with several ensemble techniques, including soft voting, maximum value, hard voting and stacking. We used three publicly available Twitter datasets (Davidson, HatEval2019, OLID) that are generated to identify offensive languages. We fused all these datasets to generate a single dataset (DHO dataset), which is more balanced across different labels, to perform multi-label classification. Our experiments have been held on Davidson dataset and the DHO corpora. The later gave the best overall results, especially F1 macro score, even it required more resources (time execution and memory). The experiments have shown good results especially the ensemble models, where stacking gave F1 score of 97% on Davidson dataset and aggregating ensembles 77% on the DHO dataset.

preprint2020arXiv

A First Instagram Dataset on COVID-19

The novel coronavirus (COVID-19) pandemic outbreak is drastically shaping and reshaping many aspects of our life, with a huge impact on our social life. In this era of lockdown policies in most of the major cities around the world, we see a huge increase in people and professional engagement in social media. Social media is playing an important role in news propagation as well as keeping people in contact. At the same time, this source is both a blessing and a curse as the coronavirus infodemic has become a major concern, and is already a topic that needs special attention and further research. In this paper, we provide a multilingual coronavirus (COVID-19) Instagram dataset that we have been continuously collected since March 30, 2020. We are making our dataset available to the research community at Github. We believe that this contribution will help the community to better understand the dynamics behind this phenomenon in Instagram, as one of the major social media. This dataset could also help study the propagation of misinformation related to this outbreak.

preprint2020arXiv

Hate Speech Detection and Racial Bias Mitigation in Social Media based on BERT model

Disparate biases associated with datasets and trained classifiers in hateful and abusive content identification tasks have raised many concerns recently. Although the problem of biased datasets on abusive language detection has been addressed more frequently, biases arising from trained classifiers have not yet been a matter of concern. Here, we first introduce a transfer learning approach for hate speech detection based on an existing pre-trained language model called BERT and evaluate the proposed model on two publicly available datasets annotated for racism, sexism, hate or offensive content on Twitter. Next, we introduce a bias alleviation mechanism in hate speech detection task to mitigate the effect of bias in training set during the fine-tuning of our pre-trained BERT-based model. Toward that end, we use an existing regularization method to reweight input samples, thereby decreasing the effects of high correlated training set' s n-grams with class labels, and then fine-tune our pre-trained BERT-based model with the new re-weighted samples. To evaluate our bias alleviation mechanism, we employ a cross-domain approach in which we use the trained classifiers on the aforementioned datasets to predict the labels of two new datasets from Twitter, AAE-aligned and White-aligned groups, which indicate tweets written in African-American English (AAE) and Standard American English (SAE) respectively. The results show the existence of systematic racial bias in trained classifiers as they tend to assign tweets written in AAE from AAE-aligned group to negative classes such as racism, sexism, hate, and offensive more often than tweets written in SAE from White-aligned. However, the racial bias in our classifiers reduces significantly after our bias alleviation mechanism is incorporated. This work could institute the first step towards debiasing hate speech and abusive language detection systems.

preprint2020arXiv

How Impersonators Exploit Instagram to Generate Fake Engagement?

Impersonators on Online Social Networks such as Instagram are playing an important role in the propagation of the content. These entities are the type of nefarious fake accounts that intend to disguise a legitimate account by making similar profiles. In addition to having impersonated profiles, we observed a considerable engagement from these entities to the published posts of verified accounts. Toward that end, we concentrate on the engagement of impersonators in terms of active and passive engagements which is studied in three major communities including ``Politician'', ``News agency'', and ``Sports star'' on Instagram. Inside each community, four verified accounts have been selected. Based on the implemented approach in our previous studies, we have collected 4.8K comments, and 2.6K likes across 566 posts created from 3.8K impersonators during 7 months. Our study shed light into this interesting phenomena and provides a surprising observation that can help us to understand better how impersonators engaging themselves inside Instagram in terms of writing Comments and leaving Likes.

preprint2016arXiv

A Trust Model for Data Sharing in Smart Cities

The data generated by the devices and existing infrastructure in the Internet of Things (IoT) should be shared among applications. However, data sharing in the IoT can only reach its full potential when multiple participants contribute their data, for example when people are able to use their smartphone sensors for this purpose. We believe that each step, from sensing the data to the actionable knowledge, requires trust-enabled mechanisms to facilitate data exchange, such as data perception trust, trustworthy data mining, and reasoning with trust related policies. The absence of trust could affect the acceptance of sharing data in smart cities. In this study, we focus on data usage transparency and accountability and propose a trust model for data sharing in smart cities, including system architecture for trust-based data sharing, data semantic and abstraction models, and a mechanism to enhance transparency and accountability for data usage. We apply semantic technology and defeasible reasoning with trust data usage policies. We built a prototype based on an air pollution monitoring use case and utilized it to evaluate the performance of our solution.

preprint2016arXiv

An EV Charging Scheduling Mechanism to Maximize User Convenience and Cost Efficiency

This paper studies charging scheduling problem of electric vehicles (EVs) in the scale of a microgrid (e.g., a university or town) where a set of charging stations are controlled by a central aggregator. A bi-objective optimization problem is formulated to jointly optimize total charging cost and user convenience. Then, a close-to-optimal online scheduling algorithm is proposed as solution. The algorithm achieves optimal charging cost and is near optimal in terms of user convenience. Moreover, the proposed method applies an efficient load forecasting technique to obtain future load information. The algorithm is assessed through simulation and compared to the previous studies. The results reveal that our method not only improves previous alternative methods in terms of Pareto-optimal solution of the bi-objective optimization problem, but also provides a close approximation for the load forecasting.

preprint2016arXiv

Maximum-Quality Tree Construction for Deadline-Constrained Aggregation in WSNs

In deadline-constrained wireless sensor networks (WSNs), quality of aggregation (QoA) is determined by the number of participating nodes in the data aggregation process. The previous studies have attempted to propose optimal scheduling algorithms to obtain the maximum QoA assuming a fixed underlying aggregation tree. However, there exists no prior work to address the issue of constructing optimal aggregation tree in deadline-constraints WSNs. The structure of underlying aggregation tree is important since our analysis demonstrates that the ratio between the maximum achievable QoAs of different trees could be as large as O(2^D), where D is the deadline. This paper casts a combinatorial optimization problem to address optimal tree construction for deadline-constrained data aggregation in WSNs. While the problem is proved to be NP-hard, we employ the recently proposed Markov approximation framework and devise two distributed algorithms with different computation overheads to find close-to-optimal solutions with bounded approximation gap. To further improve the convergence of the proposed Markov-based algorithms, we devise another initial tree construction algorithm with low computational complexity. Our extensive experiments for a set randomly-generated scenarios demonstrate that the proposed algorithms outperforms the existing alternative methods by obtaining better quality of aggregations.

preprint2016arXiv

Microgrid Revenue Maximization by Charging Scheduling of EVs in Multiple Parking Stations

Nowadays, there has been a rapid growth in global usage of the electronic vehicles (EV). Despite apparent environmental and economic advantages of EVs, their high demand charging jobs pose an immense challenge to the existing electricity grid infrastructure. In microgrids, as the small-scale version of traditional power grid, however, the EV charging scheduling is more challenging. This is because, the microgrid owner, as a large electricity customer, is interested in shaving its global peak demand, i.e., the aggregated demand over multiple parking stations, to reduce total electricity cost. While the EV charging scheduling problem in single station scenario has been studied extensively in the previous research, the microgrid-level problem with multiple stations subject to a global peak constraint is not tackled. This paper aims to propose a near-optimal EV charging scheduling mechanism in a microgrid governed by a single utility provider with multiple charging stations. The goal is to maximize the total revenue while respecting both local and global peak constraints. The underlying problem, however, is a NP-hard mixed integer linear problem which is difficult to tackle and calls for approximation algorithm design. We design a primaldual scheduling algorithm which runs in polynomial time and achieves bounded approximation ratio. Moreover, the proposed global scheduling algorithm applies a valley-filling strategy to further reduce the global peak. Simulation results show that the performance of the proposed algorithm is 98% of the optimum, which is much better than the theoretical bound obtained by our approximation analysis. Our algorithm reduces the peak demand obtained by the existing alternative algorithm by 16% and simultaneously achieves better resource utilization.

preprint2016arXiv

Wireless Sensor Network Virtualization: A Survey

Wireless Sensor Networks (WSNs) are the key components of the emerging Internet-of-Things (IoT) paradigm. They are now ubiquitous and used in a plurality of application domains. WSNs are still domain specific and usually deployed to support a specific application. However, as WSN nodes are becoming more and more powerful, it is getting more and more pertinent to research how multiple applications could share a very same WSN infrastructure. Virtualization is a technology that can potentially enable this sharing. This paper is a survey on WSN virtualization. It provides a comprehensive review of the state-of-the-art and an in-depth discussion of the research issues. We introduce the basics of WSN virtualization and motivate its pertinence with carefully selected scenarios. Existing works are presented in detail and critically evaluated using a set of requirements derived from the scenarios. The pertinent research projects are also reviewed. Several research issues are also discussed with hints on how they could be tackled.

preprint2015arXiv

A Data Annotation Architecture for Semantic Applications in Virtualized Wireless Sensor Networks

Wireless Sensor Networks (WSNs) have become very popular and are being used in many application domains (e.g. smart cities, security, gaming and agriculture). Virtualized WSNs allow the same WSN to be shared by multiple applications. Semantic applications are situation-aware and can potentially play a critical role in virtualized WSNs. However, provisioning them in such settings remains a challenge. The key reason is that semantic applications provisioning mandates data annotation. Unfortunately it is no easy task to annotate data collected in virtualized WSNs. This paper proposes a data annotation architecture for semantic applications in virtualized heterogeneous WSNs. The architecture uses overlays as the cornerstone, and we have built a prototype in the cloud environment using Google App Engine. The early performance measurements are also presented.

preprint2015arXiv

Are You Really Hidden? Predicting Current City from Profile and Social Relationship

Privacy has become a major concern in Online Social Networks (OSNs) due to threats such as advertising spam, online stalking and identity theft. Although many users hide or do not fill out their private attributes in OSNs, prior studies point out that the hidden attributes may be inferred from some other public information. Thus, users' private information could still be at stake to be exposed. Hitherto, little work helps users to assess the exposure probability/risk that the hidden attributes can be correctly predicted, let alone provides them with pointed countermeasures. In this article, we focus our study on the exposure risk assessment by a particular privacy-sensitive attribute - current city - in Facebook. Specifically, we first design a novel current city prediction approach that discloses users' hidden `current city' from their self-exposed information. Based on 371,913 Facebook users' data, we verify that our proposed prediction approach can predict users' current city more accurately than state-of-the-art approaches. Furthermore, we inspect the prediction results and model the current city exposure probability via some measurable characteristics of the self-exposed information. Finally, we construct an exposure estimator to assess the current city exposure risk for individual users, given their self-exposed information. Several case studies are presented to illustrate how to use our proposed estimator for privacy protection.

preprint2015arXiv

Reality Mining with Mobile Big Data: Understanding the Impact of Network Structure on Propagation Dynamics

Information and epidemic propagation dynamics in complex networks is truly important to discover and control terrorist attack and disease spread. How to track, recognize and model such dynamics is a big challenge. With the popularity of intellectualization and the rapid development of Internet of Things (IoT), massive mobile data is automatically collected by millions of wireless devices (e.g., smart phone and tablet). In this article, as a typical use case, the impact of network structure on epidemic propagation dynamics is investigated by using the mobile data collected from the smart phones carried by the volunteers of Ebola outbreak areas. On this basis, we propose a model to recognize the dynamic structure of a network. Then, we introduce and discuss the open issues and future work for developing the proposed recognition model.

preprint2015arXiv

Self-Modeling Based Diagnosis of Software-Defined Networks

Networks built using SDN (Software-Defined Networks) and NFV (Network Functions Virtualization) approaches are expected to face several challenges such as scalability, robustness and resiliency. In this paper, we propose a self-modeling based diagnosis to enable resilient networks in the context of SDN and NFV. We focus on solving two major problems: On the one hand, we lack today of a model or template that describes the managed elements in the context of SDN and NFV. On the other hand, the highly dynamic networks enabled by the softwarisation require the generation at runtime of a diagnosis model from which the root causes can be identified. In this paper, we propose finer granular templates that do not only model network nodes but also their sub-components for a more detailed diagnosis suitable in the SDN and NFV context. In addition, we specify and validate a self-modeling based diagnosis using Bayesian Networks. This approach differs from the state of the art in the discovery of network and service dependencies at run-time and the building of the diagnosis model of any SDN infrastructure using our templates.

preprint2015arXiv

Social and Collaborative Services for Organizations: Back to Requirements

Social and collaborative services have widely spread within the enterprises as they play a part in improving productivity and business outcomes. However, the deployment of these services fluctuates between success and failure. This paper intends to assess their deployment and how they can contribute to value creation in different industries. We investigate the relationship between the services' functionalities and the organizational requirement of these services represented by the coordination. We also consider the organizational transformation driven by servitization and emphasize its impact on the act of coordination. We highlight the tight correlation between the functionalities and the requirement in organic forms which suggests a successful deployment in such enterprises. We nonetheless find that, when the servitization is a strategic intent in organizations with mechanistic characteristics, deploying social and collaborative services can contribute to achieving this aim.

preprint2015arXiv

Wireless Sensor Network Virtualization: Early Architecture and Research Perspectives

Wireless sensor networks (WSNs) have become pervasive and are used in many applications and services. Usually the deployments of WSNs are task oriented and domain specific; thereby precluding re-use when other applications and services are contemplated. This inevitably leads to the proliferation of redundant WSN deployments. Virtualization is a technology that can aid in tackling this issue, as it enables the sharing of resources/infrastructure by multiple independent entities. In this paper we critically review the state of the art and propose a novel architecture for WSN virtualization. The proposed architecture has four layers (physical layer, virtual sensor layer, virtual sensor access layer and overlay layer) and relies on the constrained application protocol (CoAP). We illustrate its potential by using it in a scenario where a single WSN is shared by multiple applications; one of which is a fire monitoring application. We present the proof-of-concept prototype we have built along with the performance measurements, and discuss future research directions.