Source author record

Kanishka Rao

Kanishka Rao appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

7works
9topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

7 published item(s)

preprint2022arXiv

Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Large language models can encode a wealth of semantic knowledge about the world. Such knowledge could be extremely useful to robots aiming to act upon high-level, temporally extended instructions expressed in natural language. However, a significant weakness of language models is that they lack real-world experience, which makes it difficult to leverage them for decision making within a given embodiment. For example, asking a language model to describe how to clean a spill might result in a reasonable narrative, but it may not be applicable to a particular agent, such as a robot, that needs to perform this task in a particular environment. We propose to provide real-world grounding by means of pretrained skills, which are used to constrain the model to propose natural language actions that are both feasible and contextually appropriate. The robot can act as the language model's "hands and eyes," while the language model supplies high-level semantic knowledge about the task. We show how low-level skills can be combined with large language models so that the language model provides high-level knowledge about the procedures for performing complex and temporally-extended instructions, while value functions associated with these skills provide the grounding necessary to connect this knowledge to a particular physical environment. We evaluate our method on a number of real-world robotic tasks, where we show the need for real-world grounding and that this approach is capable of completing long-horizon, abstract, natural language instructions on a mobile manipulator. The project's website and the video can be found at https://say-can.github.io/.

preprint2020arXiv

RL-CycleGAN: Reinforcement Learning Aware Simulation-To-Real

Deep neural network based reinforcement learning (RL) can learn appropriate visual representations for complex tasks like vision-based robotic grasping without the need for manually engineering or prior learning a perception system. However, data for RL is collected via running an agent in the desired environment, and for applications like robotics, running a robot in the real world may be extremely costly and time consuming. Simulated training offers an appealing alternative, but ensuring that policies trained in simulation can transfer effectively into the real world requires additional machinery. Simulations may not match reality, and typically bridging the simulation-to-reality gap requires domain knowledge and task-specific engineering. We can automate this process by employing generative models to translate simulated images into realistic ones. However, this sort of translation is typically task-agnostic, in that the translated images may not preserve all features that are relevant to the task. In this paper, we introduce the RL-scene consistency loss for image translation, which ensures that the translation operation is invariant with respect to the Q-values associated with the image. This allows us to learn a task-aware translation. Incorporating this loss into unsupervised domain translation, we obtain RL-CycleGAN, a new approach for simulation-to-real-world transfer for reinforcement learning. In evaluations of RL-CycleGAN on two vision-based robotics grasping tasks, we show that RL-CycleGAN offers a substantial improvement over a number of prior methods for sim-to-real transfer, attaining excellent real-world performance with only a modest number of real-world observations.

preprint2016arXiv

Personalized Speech recognition on mobile devices

We describe a large vocabulary speech recognition system that is accurate, has low latency, and yet has a small enough memory and computational footprint to run faster than real-time on a Nexus 5 Android smartphone. We employ a quantized Long Short-Term Memory (LSTM) acoustic model trained with connectionist temporal classification (CTC) to directly predict phoneme targets, and further reduce its memory footprint using an SVD-based compression scheme. Additionally, we minimize our memory footprint by using a single language model for both dictation and voice command domains, constructed using Bayesian interpolation. Finally, in order to properly handle device-specific information, such as proper names and other context-dependent information, we inject vocabulary items into the decoder graph and bias the language model on-the-fly. Our system achieves 13.5% word error rate on an open-ended dictation task, running with a median speed that is seven times faster than real-time.

preprint2015arXiv

Fast and Accurate Recurrent Neural Network Acoustic Models for Speech Recognition

We have recently shown that deep Long Short-Term Memory (LSTM) recurrent neural networks (RNNs) outperform feed forward deep neural networks (DNNs) as acoustic models for speech recognition. More recently, we have shown that the performance of sequence trained context dependent (CD) hidden Markov model (HMM) acoustic models using such LSTM RNNs can be equaled by sequence trained phone models initialized with connectionist temporal classification (CTC). In this paper, we present techniques that further improve performance of LSTM RNN acoustic models for large vocabulary speech recognition. We show that frame stacking and reduced frame rate lead to more accurate models and faster decoding. CD phone modeling leads to further improvements. We also present initial results for LSTM RNN models outputting words directly.

preprint2012arXiv

Reinterpretion of Experimental Results with Basis Templates

Experimental analysis of data from particle collisions is typically expressed as statistical limits on a few benchmark models of particular, often historical, interest. The implications of the data for other theoretical models (current or future) may be powerful, but they cannot typically be calculated from the published information, except in the simplest case of a single-bin counting experiment. We present a novel solution to this long-standing problem by expressing the new model as a linear combination of models from published experimental analysis, allowing for the trivial calculation of limits on a nearly arbitrary model. We present tests in simple toy experiments, demonstrate self-consistency by using published results to reproduce other published results on the same spectrum, and provide a reinterpretation of a search for chiral down-type heavy quarks ($b'$) in terms of a search for an exotic heavy quark ($T$) with similar but distinct phenomenology. We find $m_T>419$ GeV at 95% CL, currently the strongest limits if the $T$ quark decays via $T\rightarrow Wb, T\rightarrow tZ$ and $T\rightarrow tH$.

preprint2012arXiv

Triangulating an exotic T quark

Limits on an exotic heavy quark $T$ are broadly generalized by considering the full range of $T\rightarrow Wb, th$ or $tZ$ branching ratios. We combine results of specific $T\rightarrow tZ$ and $T\rightarrow Wb$ searches with limits on various combinations of decay modes evaluated by re-interpreting other searches. We find strong bounds across the entire space of branching ratios, ranging from $m_T > 415$ GeV to $m_T > 557$ GeV at 95% confidence level.

preprint2012arXiv

Where are the Fermi Lines Coming From?

We estimate the spatial locations of sources of the the observed features in the Fermi-LAT photon spectrum at $E_γ=110$ and $E_γ=130$ GeV. We determine whether they are consistent with emission from a single source, as would be expected in their interpretation as $γγ$ and $γZ$ lines from dark matter annhiliation, as well as whether they are consistent with a dark matter halo positioned at the center of the galaxy. We take advantage of the per-photon measured incident angle in reconstructing the line features. In addition, we use a data-driven background model rather than making the assumption of a feature-less background. We localize the sources of the features at 110 and 130 GeV. Assuming an Einasto (NFW) density model we find the 130 GeV line to be offset from the galactic center by 285 (280) pc, the 110 GeV line by 60 (30) pc with a large relative separation of 220 (240) pc. However, we find this displacement of each source from the galactic center, as well as their relative displacement to be statistically consistent with a single Einasto or NFW dark matter halo at the center of the galaxy.