Source author record

Roberta Siciliano

Roberta Siciliano appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

5works
3topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

5 published item(s)

preprint2016arXiv

Adjusted Concordance Index, an extension of the Adjusted Rand index to fuzzy partitions

In comparing clustering partitions, Rand index (RI) and Adjusted Rand index (ARI) are commonly used for measuring the agreement between the partitions. Both these external validation indexes aim to analyze how close is a cluster to a reference (or to prior knowledge about the data) by counting corrected classified pairs of elements. When the aim is to evaluate the solution of a fuzzy clustering algorithm, the computation of these measures require converting the soft partitions into hard ones. It is known that different fuzzy partitions describing very different structures in the data can lead to the same crisp partition and consequently to the same values of these measures. We compare the existing approaches to evaluate the external validation criteria in fuzzy clustering and we propose an extension of the ARI for fuzzy partitions based on the normalized degree of concordance. Through use of real and simulated data, we analyze and evaluate the performance of our proposal.

preprint2016arXiv

Dynamic recursive tree-based partitioning for malignant melanoma identification in skin lesion dermoscopic images

In this paper, multivalued data or multiple values variables are defined. They are typical when there is some intrinsic uncertainty in data production, as the result of imprecise measuring instruments, such as in image recognition, in human judgments and so on. \noindent So far, contributions in symbolic data analysis literature provide data preprocessing criteria allowing for the use of standard methods such as factorial analysis, clustering, discriminant analysis, tree-based methods. As an alternative, this paper introduces a methodology for supervised classification, the so-called Dynamic CLASSification TREE (D-CLASS TREE), dealing simultaneously with both standard and multivalued data as well. For that, an innovative partitioning criterion with a tree-growing algorithm will be defined. Main result is a dynamic tree structure characterized by the simultaneous presence of binary and ternary partitions. A real world case study will be considered to show the advantages of the proposed methodology and main issues of the interpretation of the final results. A comparative study with other approaches dealing with the same types of data will be also shown. D-CLASS TREE outperforms its competitors in terms of accuracy, which is a fundamental aspect for predictive learning.

preprint2015arXiv

Accurate algorithms for identifying the median ranking when dealing with weak and partial rankings under the Kemeny axiomatic approach

Preference rankings virtually appear in all field of science (political sciences, behavioral sciences, machine learning, decision making and so on). The well-know social choice problem consists in trying to find a reasonable procedure to use the aggregate preferences expressed by subjects (usually called judges) to reach a collective decision. This problem turns out to be equivalent to the problem of estimating the consensus (central) ranking from data that is known to be a NP-hard Problem. Emond and Mason in 2002 proposed a branch and bound algorithm to calculate the consensus ranking given $n$ rankings expressed on $m$ objects. Depending on the complexity of the problem, there can be multiple solutions and then the consensus ranking may be not unique. We propose a new algorithm to find the consensus ranking that is equivalent to Emond and Mason's algorithm in terms of at least one of the solutions reached, but permits a really remarkable saving in computational time.

preprint2015arXiv

Boosted-Oriented Probabilistic Smoothing-Spline Clustering of Series

Fuzzy clustering methods allow the objects to belong to several clusters simultaneously, with different degrees of membership. However, a factor that influences the performance of fuzzy algorithms is the value of fuzzifier parameter. In this paper, we propose a fuzzy clustering procedure for data (time) series that does not depend on the definition of a fuzzifier parameter. It comes from two approaches, theoretically motivated for unsupervised and supervised classification cases, respectively. The first is the Probabilistic Distance (PD) clustering procedure. The second is the well known Boosting philosophy. Our idea is to adopt a boosting prospective for unsupervised learning problems, in particular we face with non hierarchical clustering problems. The aim is to assign each instance (i.e. a series) of a data set to a cluster. We assume the representative instance of a given cluster (i.e. the cluster center) as a target instance, a loss function as a synthetic index of the global performance and the probability of each instance to belong to a given cluster as the individual contribution of a given instance to the overall solution. The global performance of the proposed method is investigated by various experiments.

preprint2015arXiv

Parsimonious Time Series Clustering

We introduce a parsimonious model-based framework for clustering time course data. In these applications the computational burden becomes often an issue due to the number of available observations. The measured time series can also be very noisy and sparse and a suitable model describing them can be hard to define. We propose to model the observed measurements by using P-spline smoothers and to cluster the functional objects as summarized by the optimal spline coefficients. In principle, this idea can be adopted within all the most common clustering frameworks. In this work we discuss applications based on a k-means algorithm. We evaluate the accuracy and the efficiency of our proposal by simulations and by dealing with drosophila melanogaster gene expression data.