Source author record

Lior Shamir

Lior Shamir appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

38works
8topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

38 published item(s)

preprint2025arXiv

CornViT: A Multi-Stage Convolutional Vision Transformer Framework for Hierarchical Corn Kernel Analysis

Accurate grading of corn kernels is critical for seed certification, directional seeding, and breeding, yet it is still predominantly performed by manual inspection. This work introduces CornViT, a three-stage Convolutional Vision Transformer (CvT) framework that emulates the hierarchical reasoning of human seed analysts for single-kernel evaluation. Three sequential CvT-13 classifiers operate on 384x384 RGB images: Stage 1 distinguishes pure from impure kernels; Stage 2 categorizes pure kernels into flat and round morphologies; and Stage 3 determines the embryo orientation (up vs. down) for pure, flat kernels. Starting from a public corn seed image collection, we manually relabeled and filtered images to construct three stage-specific datasets: 7265 kernels for purity, 3859 pure kernels for morphology, and 1960 pure-flat kernels for embryo orientation, all released as benchmarks. Head-only fine-tuning of ImageNet-22k pretrained CvT-13 backbones yields test accuracies of 93.76% for purity, 94.11% for shape, and 91.12% for embryo-orientation detection. Under identical training conditions, ResNet-50 reaches only 76.56 to 81.02 percent, whereas DenseNet-121 attains 86.56 to 89.38 percent accuracy. These results highlight the advantages of convolution-augmented self-attention for kernel analysis. To facilitate adoption, we deploy CornViT in a Flask-based web application that performs stage-wise inference and exposes interpretable outputs through a browser interface. Together, the CornViT framework, curated datasets, and web application provide a deployable solution for automated corn kernel quality assessment in seed quality workflows. Source code and data are publicly available.

preprint2022arXiv

A Possible Large-scale Alignment of Galaxy Spin Directions -- Analysis of 10 Datasets from SDSS, Pan-STARRS, and HST

Multiple observations made by several different telescopes have shown asymmetry between the number of spiral galaxies rotating in opposite directions in different parts of the sky. One of the immediate questions regarding the possible asymmetry of the spin directions is whether the distribution forms a cosmological-scale axis. This paper analyzes and compares 10 different datasets published in the past decade, collected by SDSS, Pan-STARRS, and Hubble Space Telescope. The datasets contain spiral galaxies separated by their spin direction, and the distribution can show dipole axes. The analysis shows that the directions of the most probable dipole axes are consistent in datasets that have similar average redshift, but different between datasets that have different average redshift. The analysis also shows that the location of the most probable axis correlates with the average redshift of the galaxies in the datasets. That is, the location of the most probable axis shifts when the redshift gets higher, and the correlation is statistically significant. This provides a certain indication of a drift in a possible axis formed by the distribution of galaxy spin directions, or a cosmological scale structure that peaks at a certain distance from Earth.

preprint2022arXiv

Analysis of $\sim10^6$ spiral galaxies from four telescopes shows large-scale patterns of asymmetry in galaxy spin directions

The ability to collect unprecedented amounts of astronomical data has enabled the studying scientific questions that were impractical to study in the pre-information era. This study uses large datasets collected by four different robotic telescopes to profile the large-scale distribution of the spin directions of spiral galaxies. These datasets cover the Northern and Southern hemispheres, in addition to data acquired from space by the Hubble Space Telescope. The data were annotated automatically by a fully symmetric algorithm, as well as manually through a long labor-intensive process, leading to a dataset of nearly $10^6$ galaxies. The data shows possible patterns of asymmetric distribution of the spin directions, and the patterns agree between the different telescopes. The profiles also agree when using automatic or manual annotation of the galaxies, showing very similar large-scale patterns. Combining all data from all telescopes allows the most comprehensive analysis of its kind to date in terms of both the number of galaxies and the footprint size. The results show a statistically significant profile that is consistent across all telescopes. The instruments used in this study are DECam, HST, SDSS, and Pan-STARRS. The paper also discusses possible sources of bias, and analyzes the design of previous work that showed different results. Further research will be required to understand and validate these preliminary observations.

preprint2022arXiv

Analysis of spin directions of galaxies in the DESI Legacy Survey

The DESI Legacy Survey is a digital sky survey with a large footprint compared to other Earth-based surveys, covering both the Northern and Southern hemispheres. This paper shows the distribution of the spin directions of spiral galaxies imaged by DESI Legacy Survey. A simple analysis of dividing nearly 1.3$\cdot10^6$ spiral galaxies into two hemispheres shows a higher number of galaxies spinning counterclockwise in the Northern hemisphere, and a higher number of galaxies spinning clockwise in the Southern hemisphere. That distribution is consistent with previous observations, but uses a far larger number of galaxies and a larger footprint. The larger footprint allows a comprehensive analysis without the need to fit the distribution into an a priori model, making this study different from all previous analyses of this kind. Fitting the spin directions of the galaxies to cosine dependence shows a dipole axis alignment with probability of $P<10^{-5}$. The analysis is done with a trivial selection of the galaxies, as well as simple explainable annotation algorithm that does not make use of any form of machine learning, deep learning, or pattern recognition. While further work will be required, these results are aligned with previous studies suggesting the possibility of a large-scale alignment of galaxy angular momentum.

preprint2022arXiv

Asymmetry in galaxy spin directions -- analysis of data from DES and comparison to four other sky surveys

The paper shows an analysis of the large-scale distribution of galaxy spin directions of 739,286 galaxies imaged by DES. The distribution of the spin directions of the galaxies exhibits a large-scale dipole axis. Comparison of the location of the dipole axis to a similar analysis with data from SDSS, Pan-STARRS, and DESI Legacy Survey shows that all sky surveys exhibit dipole axes within 52$^o$ or less from each other, well within 1$σ$ error. While non-random distribution is unexpected, the findings are consistent across all sky surveys, regardless of the telescope or whether the data were annotated manually or automatically. Possible errors that can lead to the observation are discussed. The paper also discusses previous studies showing opposite conclusions,and analyzes the decisions that led to these results. Although the observation is provocative, and further research will be required, the existing evidence justifies to consider the contention that galaxy spin directions as observed from Earth are not necessarily randomly distributed. Possible explanations can be related to mature cosmological theories, but also to the internal structure of galaxies.

preprint2022arXiv

Large-scale asymmetry in galaxy spin directions -- analysis of galaxies with spectra in DES, SDSS, and DESI Legacy Survey

Multiple previous studies using several different probes have shown considerable evidence for the existence of cosmological-scale anisotropy and a Hubble-scale axis. One of the probes that show such evidence is the distribution of the directions toward which galaxies spin. The advantage of the analysis of the distribution of galaxy spin directions compared to the CMB anisotropy is that the ratio of galaxy spin directions is a relative measurement, and therefore less sensitive to background contamination such as Milky Way obstruction. Another advantage is that many spiral galaxies have spectra, and therefore allow to analyze the location of such axis relative to Earth. This paper shows an analysis of the distribution of the spin directions of over 90K galaxies with spectra. That analysis is also compared to previous analyses using the Earth-based SDSS, Pan-STARRS, and DESI Legacy Survey, as well as space-based data collected by HST. The results show very good agreement between the distribution patterns observed with the different telescopes. The dipole or quadrupole axes formed by the spin directions of the galaxies with spectra do not necessarily go directly through Earth.

preprint2022arXiv

New evidence and analysis of cosmological-scale asymmetry in galaxy spin directions

In the past several decades, multiple cosmological theories that are based on the contention that the Universe has a major axis have been proposed. Such theories can be based on the geometry of the Universe, or multiverse theories such as black hole cosmology. The contention of a cosmological-scale axis is supported by certain evidence such as the dipole axis formed by the CMB distribution. Here I study another form of cosmological-scale axis, based on the distribution of the spin direction of spiral galaxies. Data from four different telescopes is analyzed, showing nearly identical axis profiles when the distribution of the redshifts of the galaxies is similar.

preprint2022arXiv

Self-Supervised Approach to Addressing Zero-Shot Learning Problem

In recent years, self-supervised learning has had significant success in applications involving computer vision and natural language processing. The type of pretext task is important to this boost in performance. One common pretext task is the measure of similarity and dissimilarity between pairs of images. In this scenario, the two images that make up the negative pair are visibly different to humans. However, in entomology, species are nearly indistinguishable and thus hard to differentiate. In this study, we explored the performance of a Siamese neural network using contrastive loss by learning to push apart embeddings of bumblebee species pair that are dissimilar, and pull together similar embeddings. Our experimental results show a 61% F1-score on zero-shot instances, a performance showing 11% improvement on samples of classes that share intersections with the training set.

preprint2022arXiv

Systematic biases when using deep neural networks for annotating large catalogs of astronomical images

Deep convolutional neural networks (DCNNs) have become the most common solution for automatic image annotation due to their non-parametric nature, good performance, and their accessibility through libraries such as TensorFlow. Among other fields, DCNNs are also a common approach to the annotation of large astronomical image databases acquired by digital sky surveys. One of the main downsides of DCNNs is the complex non-intuitive rules that make DCNNs act as a ``black box", providing annotations in a manner that is unclear to the user. Therefore, the user is often not able to know what information is used by the DCNNs for the classification. Here we demonstrate that the training of a DCNN is sensitive to the context of the training data such as the location of the objects in the sky. We show that for basic classification of elliptical and spiral galaxies, the sky location of the galaxies used for training affects the behavior of the algorithm, and leads to a small but consistent and statistically significant bias. That bias exhibits itself in the form of cosmological-scale anisotropy in the distribution of basic galaxy morphology. Therefore, while DCNNs are powerful tools for annotating images of extended sources, the construction of training sets for galaxy morphology should take into consideration more aspects than the visual appearance of the object. In any case, catalogs created with deep neural networks that exhibit signs of cosmological anisotropy should be interpreted with the possibility of consistent bias.

preprint2021arXiv

AdeNet: Deep learning architecture that identifies damaged electrical insulators in power lines

Ceramic insulators are important to electronic systems, designed and installed to protect humans from the danger of high voltage electric current. However, insulators are not immortal, and natural deterioration can gradually damage them. Therefore, the condition of insulators must be continually monitored, which is normally done using UAVs. UAVs collect many images of insulators, and these images are then analyzed to identify those that are damaged. Here we describe AdeNet as a deep neural network designed to identify damaged insulators, and test multiple approaches to automatic analysis of the condition of insulators. Several deep neural networks were tested, as were shallow learning methods. The best results (88.8\%) were achieved using AdeNet without transfer learning. AdeNet also reduced the false negative rate to $\sim$7\%. While the method cannot fully replace human inspection, its high throughput can reduce the amount of labor required to monitor lines for damaged insulators and provide early warning to replace damaged insulators.

preprint2021arXiv

Analysis of the alignment of non-random patterns of spin directions in populations of spiral galaxies

Observations of non-random distribution of galaxies with opposite spin directions have recently attracted considerable attention. Here, a method for identifying cosine-dependence in a dataset of galaxies annotated by their spin directions is described in the light of different aspects that can impact the statistical analysis of the data. These aspects include the presence of duplicate objects in a dataset, errors in the galaxy annotation process, and non-random distribution of the asymmetry that does not necessarily form a dipole or quadrupole axes. The results show that duplicate objects in the dataset can artificially increase the likelihood of cosine dependence detected in the data, but a very high number of duplicate objects is required to lead to a false detection of an axis. Inaccuracy in galaxy annotations has relatively minor impact on the identification of cosine dependence when the error is randomly distributed between clockwise and counterclockwise galaxies. However, when the error is not random, even a small bias of 1% leads to a statistically significant cosine dependence that peaks at the celestial pole. Experiments with artificial datasets in which the distribution was not random showed strong cosine dependence even when the data did not form a full dipole axis alignment. The analysis when using the unmodified data shows asymmetry profile similar to the profile shown in multiple previous studies using several different telescopes.

preprint2021arXiv

Automatic identification of outliers in Hubble Space Telescope galaxy images

Rare extragalactic objects can carry substantial information about the past, present, and future universe. Given the size of astronomical databases in the information era it can be assumed that very many outlier galaxies are included in existing and future astronomical databases. However, manual search for these objects is impractical due to the required labor, and therefore the ability to detect such objects largely depends on computer algorithms. This paper describes an unsupervised machine learning algorithm for automatic detection of outlier galaxy images, and its application to several Hubble Space Telescope fields. The algorithm does not require training, and therefore is not dependent on the preparation of clean training sets. The application of the algorithm to a large collection of galaxies detected a variety of outlier galaxy images. The algorithm is not perfect in the sense that not all objects detected by the algorithm are indeed considered outliers, but it reduces the dataset by two orders of magnitude to allow practical manual identification. The catalogue contains 147 objects that would be very difficult to identify without using automation.

preprint2020arXiv

Asymmetry between galaxies with different spin patterns: A comparison between COSMOS, SDSS, and Pan-STARRS

Previous observations of a large number of galaxies show differences between the photometry of spiral galaxies with clockwise spin patterns and spiral galaxies with counterclockwise spin patterns. In this study the mean magnitude of a large number of clockwise galaxies is compared to the mean magnitude of a large number of counterclockwise galaxies. The observed difference between clockwise and counterclockwise spiral galaxies imaged by the space-based COSMOS survey is compared to the differences between clockwise and counterclockwise galaxies imaged by the Earth-based SDSS and Pan-STARRS around the same field. The annotation of clockwise and counterclockwise galaxies is a fully automatic process that does not involve human intervention, and in all experiments both clockwise and counterclockwise galaxies are separated from the same fields. The comparison shows that the same asymmetry was identified by all three telescopes, providing strong evidence that the rotation direction of a spiral galaxy is linked to its luminosity as measured from Earth. Analysis of the luminosity difference using a large number of galaxies from different parts of the sky shows that the difference between clockwise and counterclockwise galaxies changes with the direction of observation, and oriented around an axis.

preprint2020arXiv

Eliminating self-selection: Using data science for authentic undergraduate research in a first-year introductory course

Research experience and mentoring has been identified as an effective intervention for increasing student engagement and retention in the STEM fields, with high impact on students from undeserved populations. However, one-on-one mentoring is limited by the number of available faculty, and in certain cases also by the availability of funding for stipend. One-on-one mentoring is further limited by the selection and self-selection of students. Since research positions are often competitive, they are often taken by the best-performing students. More importantly, many students who do not see themselves as the top students of their class, or do not identify themselves as researchers might not apply, and that self selection can have the highest impact on non-traditional students. To address the obstacles of scalability, selection, and self-selection, we designed a data science research experience for undergraduates as part of an introductory computer science course. Through the intervention, the students are exposed to authentic research as early as their first semester. The intervention is inclusive in the sense that all students registered to the course participate in the research, with no process of selection or self-selection. The research is focused on analytics of large text databases. Using discovery-enabling software tools, the students analyze a corpus of congressional speeches, and identify patterns of differences between democratic speeches and republican speeches, differences between speeches for and against certain bills, and differences between speeches about bills that passed and bills that did not pass. In the beginning of the research experience all student follow the same protocol and use the same data, and then each group of students work on their own research project as part of their final project of the course.

preprint2020arXiv

Large-scale asymmetry between clockwise and counterclockwise galaxies revisited

The ability of digital sky surveys to collect and store very large amounts of data provides completely new ways to study the local universe. Perhaps one of the most provocative observations reported with such tools is the asymmetry between galaxies with clockwise and counterclockwise spin patterns. Here I use $\sim1.7\cdot10^5$ spiral galaxies from SDSS and sort them by their spin patterns (clockwise or counterclockwise) to identify and profile a possible large-scale pattern of the distribution of galaxy spin patterns as observed from Earth. The analysis shows asymmetry between the number of clockwise and counterclockwise spiral galaxies imaged by SDSS, and a dipole axis. These findings largely agree with previous reports using smaller datasets. The probability of the differences between the number of galaxies to occur by chance is (P<4*10^-9), and the probability of an asymmetry axis to occur by mere chance is (P<1.4*10^-5).

preprint2020arXiv

Patterns of galaxy spin directions in SDSS and Pan-STARRS show parity violation and multipoles

The distribution of spin directions of $\sim6.4\cdot10^4$ SDSS spiral galaxies with spectra was examined, and compared to the distribution of $\sim3.3\cdot10^4$ Pan-STARRS galaxies. The analysis shows a statistically significant asymmetry between the number of SDSS galaxies with opposite spin directions, and the magnitude and direction of the asymmetry changes with the direction of observation and with the redshift. The redshift dependence shows that the distribution of the spin direction of SDSS galaxies becomes more asymmetric as the redshift gets higher. Fitting the distribution of the galaxy spin directions to a quadrupole alignment provides fitness with statistical significance >5$σ$, which grows to >8$σ$ when just galaxies with z>0.15 are used. Similar analysis with Pan-STARRS galaxies provides dipole and quadrupole alignments nearly identical to the analysis of SDSS galaxies, showing that the source of the asymmetry is not necessarily a certain unknown flaw in a specific telescope system. While these observations are clearly provocative, there is no known error that could exhibit itself in such form. The data analysis process is fully automatic, and uses deterministic and symmetric algorithms with defined rules. It does not involve either manual analysis that can lead to human perceptual bias, or machine learning that can capture human biases or other subtle differences that are difficult to identify due to the complex nature of machine learning processes. Also, an error in the galaxy annotation process is expected to show consistent bias in all parts of the sky, rather than change with the direction of observation to form a clear and definable pattern.

preprint2016arXiv

Astrophysics Source Code Library: Here we grow again!

The Astrophysics Source Code Library (ASCL) is a free online registry of research codes; it is indexed by ADS and Web of Science and has over 1300 code entries. Its entries are increasingly used to cite software; citations have been doubling each year since 2012 and every major astronomy journal accepts citations to the ASCL. Codes in the resource cover all aspects of astrophysics research and many programming languages are represented. In the past year, the ASCL added dashboards for users and administrators, started minting Digital Objective Identifiers (DOIs) for software it houses, and added metadata fields requested by users. This presentation covers the ASCL's growth in the past year and the opportunities afforded it as one of the few domain libraries for science research codes.

preprint2016arXiv

Asymmetry between galaxies with clockwise handedness and counterclockwise handedness

While it is clear that spiral galaxies can have different handedness, galaxies with clockwise patterns are assumed to be symmetric in all of their other characteristics to galaxies with counterclockwise patterns. Here we use data from SDSS DR7 to show that photometric data can distinguish between clockwise and counterclockwise galaxies. Pattern recognition algorithms trained and tested using the photometric data of a clean manually crafted dataset of 13,440 spiral galaxies with z<0.25 can predict the handedness of a spiral galaxy in ~64% of the cases, significantly higher than mere chance accuracy of 50% (P<10^{-5}). Experiments with a different dataset of 10,281 automatically classified galaxies showed similar results of $~65% classification accuracy, suggesting that the observed asymmetry is consistent also in datasets annotated in a fully automatic process, and without human intervention. That shows that the photometric data collected by SDSS is sensitive to the handedness of the galaxy. Also, analysis of the number of galaxies classified as clockwise and counterclockwise by crowdsourcing shows that manual classification between spiral and elliptical galaxies can be affected by the handedness of the galaxy, and therefore galaxy morphology analyzed by citizen science campaigns might be biased by the galaxy handedness. Code and data used in the experiment are publicly available, and the experiment can be easily replicated.

preprint2016arXiv

Computer-generated visual morphology catalog of ~3,000,000 SDSS galaxies

We applied computer analysis to classify the broad morphological type of ~3,000,000 SDSS galaxies. The catalog provides for each galaxy the DR8 object ID, right ascension, declination, and the certainty of the automatic classification to spiral or elliptical. The certainty of the classification allows controlling the accuracy of a subset of galaxies by sacrificing some of the least certain classifications. The accuracy of the catalog was tested using galaxies that were classified by the manually annotated Galaxy Zoo catalog. The results show that the catalog contains ~900,000 spiral galaxies and ~600,000 elliptical galaxies with classification certainty that has a statistical agreement rate of ~98% with Galaxy Zoo debiased 'superclean' dataset. That also demonstrates the ability of computers to turn large datasets of galaxy images into structured catalogs of galaxy morphology. The catalog can be downloaded at http://vfacstaff.ltu.edu/lshamir/data/morph_catalog , and can be accessed through public tables on CAS: public.broadMorph.LargeGM, public.broadMorph.LargeWnnGM, and public.broadMorph.SpectraGM. The image analysis software that was used to create the catalog is also publicly available.

preprint2016arXiv

Morphology-based query for galaxy image databases

Galaxies of rare morphology are of paramount scientific interest, as they carry important information about the past, present, and future universe. Once a rare galaxy is identified, studying it more effectively requires a set of galaxies of similar morphology, allowing generalization and statistical analysis that cannot be done when $N=1$. Databases generated by digital sky surveys can contain a very large number of galaxy images, and therefore once a rare galaxy of interest is identified it is possible that more instances of the same morphology are also present in the database. However, when a researcher identifies a certain galaxy of rare morphology in the database, it is virtually impossible to mine the database manually in the search for galaxies of similar morphology. Here we propose a computer method that can automatically search databases of galaxy images and identify galaxies that are morphologically similar to a certain user-defined query galaxy. That is, the researcher provides an image of a galaxy of interest, and the pattern recognition system automatically returns a list of galaxies that are visually similar to the target galaxy. The algorithm uses a comprehensive set of descriptors, allowing it to support different types of galaxies, and it is not limited to a finite set of known morphology. While the list of returned galaxies is neither clean nor complete, it contains a far higher frequency of galaxies of the morphology of interest, providing a substantial reduction of the data. Such algorithms can be integrated into data management systems of autonomous digital sky surveys such as the Large Synoptic Survey Telescope (LSST), where the number of galaxies in the database is extremely large. The source code of the method is available at http://vfacstaff.ltu.edu/lshamir/downloads/udat.

preprint2015arXiv

Galaxy morphology - an unsupervised machine learning approach

Structural properties posses valuable information about the formation and evolution of galaxies, and are important for understanding the past, present, and future universe. Here we use unsupervised machine learning methodology to analyze a network of similarities between galaxy morphological types, and automatically deduce a morphological sequence of galaxies. Application of the method to the EFIGI catalog show that the morphological scheme produced by the algorithm is largely in agreement with the De Vaucouleurs system, demonstrating the ability of computer vision and machine learning methods to automatically profile galaxy morphological sequences. The unsupervised analysis method is based on comprehensive computer vision techniques that compute the visual similarities between the different morphological types. Rather than relying on human cognition, the proposed system deduces the similarities between sets of galaxy images in an automatic manner, and is therefore not limited by the number of galaxies being analyzed. The source code of the method is publicly available, and the protocol of the experiment is included in the paper so that the experiment can be replicated, and the method can be used to analyze user-defined datasets of galaxy images.

preprint2015arXiv

Improving Software Citation and Credit

The past year has seen movement on several fronts for improving software citation, including the Center for Open Science's Transparency and Openness Promotion (TOP) Guidelines, the Software Publishing Special Interest Group that was started at January's AAS meeting in Seattle at the request of that organization's Working Group on Astronomical Software, a Sloan-sponsored meeting at GitHub in San Francisco to begin work on a cohesive research software citation-enabling platform, the work of Force11 to "transform and improve" research communication, and WSSSPE's ongoing efforts that include software publication, citation, credit, and sustainability. Brief reports on these efforts were shared at the BoF, after which participants discussed ideas for improving software citation, generating a list of recommendations to the community of software authors, journal publishers, ADS, and research authors. The discussion, recommendations, and feedback will help form recommendations for software citation to those publishers represented in the Software Publishing Special Interest Group and the broader community.

preprint2014arXiv

Astrophysics Source Code Library Enhancements

The Astrophysics Source Code Library (ASCL; ascl.net) is a free online registry of codes used in astronomy research; it currently contains over 900 codes and is indexed by ADS. The ASCL has recently moved a new infrastructure into production. The new site provides a true database for the code entries and integrates the WordPress news and information pages and the discussion forum into one site. Previous capabilities are retained and permalinks to ascl.net continue to work. This improvement offers more functionality and flexibility than the previous site, is easier to maintain, and offers new possibilities for collaboration. This presentation covers these recent changes to the ASCL.

preprint2014arXiv

Automatic detection of peculiar galaxy pairs in Sloan Digital Sky Survey

We applied computational tools for automatic detection of peculiar galaxy pairs. We first detected in SDSS DR7 ~400,000 galaxy images with i magnitude <18 that had more than one point spread function, and then applied a machine learning algorithm that detected ~26,000 galaxy images that had morphology similar to the morphology of galaxy mergers. That dataset was mined using a novelty detection algorithm, producing a short list of 500 most peculiar galaxies as quantitatively determined by the algorithm. Manual examination of these galaxies showed that while most of the galaxy pairs in the list were not necessarily peculiar, numerous unusual galaxy pairs were detected. In this paper we describe the protocol and computational tools used for the detection of peculiar mergers, and provide examples of peculiar galaxy pairs that were detected.

preprint2014arXiv

Combining human and machine learning for morphological analysis of galaxy images

The increasing importance of digital sky surveys collecting many millions of galaxy images has reinforced the need for robust methods that can perform morphological analysis of large galaxy image databases. Citizen science initiatives such as Galaxy Zoo showed that large datasets of galaxy images can be analyzed effectively by non-scientist volunteers, but since databases generated by robotic telescopes grow much faster than the processing power of any group of citizen scientists, it is clear that computer analysis is required. Here we propose to use citizen science data for training machine learning systems, and show experimental results demonstrating that machine learning systems can be trained with citizen science data. Our findings show that the performance of machine learning depends on the quality of the data, which can be improved by using samples that have a high degree of agreement between the citizen scientists. The source code of the method is publicly available.

preprint2013arXiv

Astrophysics Source Code Library: Incite to Cite!

The Astrophysics Source Code Library (ASCL, http://ascl.net/) is an online registry of over 700 source codes that are of interest to astrophysicists, with more being added regularly. The ASCL actively seeks out codes as well as accepting submissions from the code authors, and all entries are citable and indexed by ADS. All codes have been used to generate results published in or submitted to a refereed journal and are available either via a download site or froman identified source. In addition to being the largest directory of scientist-written astrophysics programs available, the ASCL is also an active participant in the reproducible research movement with presentations at various conferences, numerous blog posts and a journal article. This poster provides a description of the ASCL and the changes that we are starting to see in the astrophysics community as a result of the work we are doing.

preprint2013arXiv

Automatic quantitative morphological analysis of interacting galaxies

The large number of galaxies imaged by digital sky surveys reinforces the need for computational methods for analyzing galaxy morphology. While the morphology of most galaxies can be associated with a stage on the Hubble sequence, morphology of galaxy mergers is far more complex due to the combination of two or more galaxies with different morphologies and the interaction between them. Here we propose a computational method based on unsupervised machine learning that can quantitatively analyze morphologies of galaxy mergers and associate galaxies by their morphology. The method works by first generating multiple synthetic galaxy models for each galaxy merger, and then extracting a large set of numerical image content descriptors for each galaxy model. These numbers are weighted using Fisher discriminant scores, and then the similarities between the galaxy mergers are deduced using a variation of Weighted Nearest Neighbor analysis such that the Fisher scores are used as weights. The similarities between the galaxy mergers are visualized using phylogenies to provide a graph that reflects the morphological similarities between the different galaxy mergers, and thus quantitatively profile the morphology of galaxy mergers.

preprint2013arXiv

Color Differences between Clockwise and Counterclockwise Spiral Galaxies

While spiral galaxies observed from Earth clearly seem to spin in different directions, little is yet known about other differences between galaxies that spin clockwise and galaxies that spin counterclockwise. Here we compared the color of 64,399 spiral galaxies that spin clockwise to 63,215 spiral galaxies that spin counterclockwise. The results show that clockwise galaxies tend to be bluer than galaxies that spin counterclockwise. The probability that the color differences can be attributed to chance is ~0.019.

preprint2013arXiv

Ideas for Advancing Code Sharing (A Different Kind of Hack Day)

How do we as a community encourage the reuse of software for telescope operations, data processing, and calibration? How can we support making codes used in research available for others to examine? Continuing the discussion from last year Bring out your codes! BoF session, participants separated into groups to brainstorm ideas to mitigate factors which inhibit code sharing and nurture those which encourage code sharing. The BoF concluded with the sharing of ideas that arose from the brainstorming sessions and a brief summary by the moderator.

preprint2013arXiv

Practices in source code sharing in astrophysics

While software and algorithms have become increasingly important in astronomy, the majority of authors who publish computational astronomy research do not share the source code they develop, making it difficult to replicate and reuse the work. In this paper we discuss the importance of sharing scientific source code with the entire astrophysics community, and propose that journals require authors to make their code publicly available when a paper is published. That is, we suggest that a paper that involves a computer program not be accepted for publication unless the source code becomes publicly available. The adoption of such a policy by editors, editorial boards, and reviewers will improve the ability to replicate scientific results, and will also make the computational astronomy methods more available to other researchers who wish to apply them to their data.

preprint2013arXiv

Quantitative analysis of spirality in elliptical galaxies

We use an automated galaxy morphology analysis method to quantitatively measure the spirality of galaxies classified manually as elliptical. The data set used for the analysis consists of 60,518 galaxy images with redshift obtained by the Sloan Digital Sky Survey (SDSS) and classified manually by Galaxy Zoo, as well as the RC3 and NA10 catalogues. We measure the spirality of the galaxies by using the Ganalyzer method, which transforms the galaxy image to its radial intensity plot to detect galaxy spirality that is in many cases difficult to notice by manual observation of the raw galaxy image. Experimental results using manually classified elliptical and S0 galaxies with redshift <0.3 suggest that galaxies classified manually as elliptical and S0 exhibit a nonzero signal for the spirality. These results suggest that the human eye observing the raw galaxy image might not always be the most effective way of detecting spirality and curves in the arms of galaxies.

preprint2013arXiv

The Astrophysics Source Code Library: Where do we go from here?

The Astrophysics Source Code Library, started in 1999, has in the past three years grown from a repository for 40 codes to a registry of over 700 codes that are now indexed by ADS. What comes next? We examine the future of the ASCL, the challenges facing it, the rationale behind its practices, and the need to balance what we might do with what we have the resources to accomplish.

preprint2012arXiv

Handedness asymmetry of spiral galaxies with z<0.3 shows cosmic parity violation and a dipole axis

A dataset of 126,501 spiral galaxies taken from Sloan Digital Sky Survey was used to analyze the large-scale galaxy handedness in different regions of the local universe. The analysis was automated by using a transformation of the galaxy images to their radial intensity plots, which allows automatic analysis of the galaxy spin and can therefore be used to analyze a large galaxy dataset. The results show that the local universe (z<0.3) is not isotropic in terms of galaxy spin, with probability P<5.8*10^-6 of such asymmetry to occur by chance. The handedness asymmetries exhibit an approximate cosine dependence, and the most likely dipole axis was found at RA=132, DEC=32 with 1 sigma error range of 107 to 179 degrees for the RA. The probability of such axis to occur by chance is P<1.95*10^-5 . The amplitude of the handedness asymmetry reported in this paper is generally in agreement with Longo, but the statistical significance is improved by a factor of 40, and the direction of the axis disagrees somewhat.

preprint2012arXiv

Practices in Code Discoverability

Much of scientific progress now hinges on the reliability, falsifiability and reproducibility of computer source codes. Astrophysics in particular is a discipline that today leads other sciences in making useful scientific components freely available online, including data, abstracts, preprints, and fully published papers, yet even today many astrophysics source codes remain hidden from public view. We review the importance and history of source codes in astrophysics and previous efforts to develop ways in which information about astrophysics codes can be shared. We also discuss why some scientist coders resist sharing or publishing their codes, the reasons for and importance of overcoming this resistance, and alert the community to a reworking of one of the first attempts for sharing codes, the Astrophysics Source Code Library (ASCL). We discuss the implementation of the ASCL in an accompanying poster paper. We suggest that code could be given a similar level of referencing as data gets in repositories such as ADS.

preprint2012arXiv

Practices in Code Discoverability: Astrophysics Source Code Library

Here we describe the Astrophysics Source Code Library (ASCL), which takes an active approach to sharing astrophysical source code. ASCL's editor seeks out both new and old peer-reviewed papers that describe methods or experiments that involve the development or use of source code, and adds entries for the found codes to the library. This approach ensures that source codes are added without requiring authors to actively submit them, resulting in a comprehensive listing that covers a significant number of the astrophysics source codes used in peer-reviewed studies. The ASCL now has over 340 codes in it and continues to grow. In 2011, the ASCL (http://ascl.net) has on average added 19 new codes per month. An advisory committee has been established to provide input and guide the development and expansion of the new site, and a marketing plan has been developed and is being executed. All ASCL source codes have been used to generate results published in or submitted to a refereed journal and are freely available either via a download site or from an identified source. This paper provides the history and description of the ASCL. It lists the requirements for including codes, examines the benefits of the ASCL, and outlines some of its future plans.

preprint2011arXiv

Ganalyzer: A tool for automatic galaxy image analysis

We describe Ganalyzer, a model-based tool that can automatically analyze and classify galaxy images. Ganalyzer works by separating the galaxy pixels from the background pixels, finding the center and radius of the galaxy, generating the radial intensity plot, and then computing the slopes of the peaks detected in the radial intensity plot to measure the spirality of the galaxy and determine its morphological class. Unlike algorithms that are based on machine learning, Ganalyzer is based on measuring the spirality of the galaxy, a task that is difficult to perform manually, and in many cases can provide a more accurate analysis compared to manual observation. Ganalyzer is simple to use, and can be easily embedded into other image analysis applications. Another advantage is its speed, which allows it to analyze ~10,000,000 galaxy images in five days using a standard modern desktop computer. These capabilities can make Ganalyzer a useful tool in analyzing large datasets of galaxy images collected by autonomous sky surveys such as SDSS, LSST or DES. The software is available for free download at http://vfacstaff.ltu.edu/lshamir/downloads/ganalyzer, and the data used in the experiment are available at http://vfacstaff.ltu.edu/lshamir/downloads/ganalyzer/GalaxyImages.zip.

preprint2009arXiv

Automatic morphological classification of galaxy images

We describe an image analysis supervised learning algorithm that can automatically classify galaxy images. The algorithm is first trained using a manually classified images of elliptical, spiral, and edge-on galaxies. A large set of image features is extracted from each image, and the most informative features are selected using Fisher scores. Test images can then be classified using a simple Weighted Nearest Neighbor rule such that the Fisher scores are used as the feature weights. Experimental results show that galaxy images from Galaxy Zoo can be classified automatically to spiral, elliptical and edge-on galaxies with accuracy of ~90% compared to classifications carried out by the author. Full compilable source code of the algorithm is available for free download, and its general-purpose nature makes it suitable for other uses that involve automatic image analysis of celestial objects.

preprint2009arXiv

Frequency Limits on Naked-Eye Optical Transients Lasting from Minutes to Years

How often do bright optical transients occur on the sky but go unreported? To constrain the bright end of the astronomical transient function, a systematic search for transients that become bright enough to be noticed by the unaided eye was conducted using the all-sky monitors of the Night Sky Live network. Two fisheye continuous cameras (CONCAMs) operating over three years created a data base that was searched for transients that appeared in time-contiguous CCD frames. Although a single candidate transient was found (Nemiroff and Shamir 2006), the lack of more transients is used here to deduce upper limits to the general frequency of bright transients. To be detected, a transient must have increased by over three visual magnitudes to become brighter than visual magnitude 5.5 on the time scale of minutes to years. It is concluded that, on the average, fewer than 0.0040 ($t_{dur} / 60$ seconds) transients with duration $t_{dur}$ between minutes and hours, occur anywhere on the sky at any one time. For transients on the order of months to years, fewer than 160 ($t_{dur} / 1$ year) occur, while for transients on the order of years to millennia, fewer than 50 ($t_{dur}/1$ year)$^2$ occur.