arXiv ScienceSearch

arXiv subjects

Dibyendu Ghosh

Publications and source records attributed to Dibyendu Ghosh.

8 recordsLinked to original sources

OVMAN: A Task and Benchmark for Open-Vocabulary Motion-Aware Navigation

Homes change between a robot's visits. Navigation benchmarks pose their goals in the world the agent currently sees, and the two-visit benchmarks that exist score recall or rearrangement rather than navigation. None of them can express go to the chair that was moved or go to where the vase used to be. OVMAN is a task in which an agent tours a scene, returns after a scripted change, and must navigate to a goal specified by the change itself. Two of its six change relations answer with a place an object has left, where nothing remains to be detected. We release 219 two-visit episodes, each certified solvable by an oracle agent that completes it three times. Two released systems fail as predicted. A zero-shot object-goal navigator reaches the change-defined target in 12.3% of episodes and almost never reaches a vacated location, 0.026 on former and 0.050 on removed. When a self-maintaining open-vocabulary map is read as two visits rather than as one maintained map, overall success rises (0.396 to 0.479 on our mapping) but the past-position relations specifically do not recover, on our maps or on a released self-maintaining one; keeping a map current is therefore not the whole obstacle. A simple two-visit reference agent reaches 45.2% when navigating, against an embodied oracle of 99.5%. An error decomposition places the remaining difficulty in carrying an instance identity across visits rather than in naming it or in selecting the answer once positions are known.

cs.RO

Memory for Attention: Language-Conditioned Re-Perception with a Vision--Language--Motion Map

A robot carrying a persistent, behavior-annotated map faces two very different planning questions, and its memory answers only one of them well. The spatial-navigation question - how to walk around a room - we address first, and report a negative: building on Vision-Language-Motion Maps (VLMM), a behavior-aware planner cost cuts a planning-time objective by ~35% over 28 AI2-THOR scenes, but under closed-loop execution the benefit nearly vanishes (~4%) and an on-demand vision-language model (VLM) does as well. The resource-allocation question is different: under a limited perception budget, what should the robot attend to right now to keep its own map fresh? Framing re-perception as this attention decision, we show a persistent map's memory - change-history, or even just recency of last sighting - yields the best of the re-perception schedules we test (held-out), while the memoryless category prior is the weakest of them under the sqrt-law schedule -- though under a Whittle index it leads memory until heterogeneity is real. The gain grows with per-instance heterogeneity as a Cauchy-Schwarz bound predicts, tracking Var(sqrt(lambda)), the variance of root-volatility, and reallocates budget toward the important objects the schedule protects; against a real CLIP movability prior it is +21-26%, of which roughly half survives once that prior's saturated scale is calibrated (+7-13%). The map's full combination earns its keep when the task is language-conditioned: told what to keep track of, VLMM grounds the relevant objects (open-vocabulary) and tracks their change (memory), beating a relevance-weighted recency baseline (+2.9% over 26 queries, at full heterogeneity; the ordering reverses when instance rates track category norms) - and a category prior (+9.2%). The map earns its keep not by telling the robot how to walk around a room, but by telling it what to pay attention to.

cs.RO

Vision-Language-Motion Maps: An Open-Vocabulary, Uncertainty-Aware, Queryable Motion Attribute for 3D Scene Maps

Open-vocabulary 3D maps let robots answer language queries about what and where, but they assume a static world and cannot answer queries about how scene elements behave. We introduce Vision-Language-Motion Maps (VLMM), an open-vocabulary, language-queryable 3D map - queried through a rule-based intent router over open-vocabulary object nouns, not a general natural-language interface - in which each element carries a fused motion attribute: a VLM/LLM semantic movability prior combined with geometrically observed cross-frame motion, together with a per-element uncertainty. Queries reduce to attribute filters that distinguish what has been seen to move, what could move but has not, and what stays still. On a controlled simulator benchmark with exact ground truth (AI2-THOR, three scene types) we show through ablation that the schema fields are non-substitutable: a semantic-only baseline fails motion queries even with strong features, and neither motion field substitutes for the other (the prior cannot answer "what is moving," observed motion cannot answer "what could move"). On real dynamic RGB-D (TUM and Bonn, six sequences) we show the uncertainty channel - our key difference from prior fused-motion work - consistently improves moving-vs-static average precision and reduces false motion flags, and that it is robust to estimated (noisy) poses. The raw confidence is not calibrated, but post-hoc isotonic calibration reaches an expected calibration error of 0.10. VLMM is a representation contribution: the closest prior maps each lack at least one of the four properties - open-vocabulary, language-queryable, fused prior-and-observed motion, and per-element uncertainty - that our combination provides.

cs.RO

ICPR 2024 Competition on Safe Segmentation of Drive Scenes in Unstructured Traffic and Adverse Weather Conditions

The ICPR 2024 Competition on Safe Segmentation of Drive Scenes in Unstructured Traffic and Adverse Weather Conditions served as a rigorous platform to evaluate and benchmark state-of-the-art semantic segmentation models under challenging conditions for autonomous driving. Over several months, participants were provided with the IDD-AW dataset, consisting of 5000 high-quality RGB-NIR image pairs, each annotated at the pixel level and captured under adverse weather conditions such as rain, fog, low light, and snow. A key aspect of the competition was the use and improvement of the Safe mean Intersection over Union (Safe mIoU) metric, designed to penalize unsafe incorrect predictions that could be overlooked by traditional mIoU. This innovative metric emphasized the importance of safety in developing autonomous driving systems. The competition showed significant advancements in the field, with participants demonstrating models that excelled in semantic segmentation and prioritized safety and robustness in unstructured and adverse conditions. The results of the competition set new benchmarks in the domain, highlighting the critical role of safety in deploying autonomous vehicles in real-world scenarios. The contributions from this competition are expected to drive further innovation in autonomous driving technology, addressing the critical challenges of operating in diverse and unpredictable environments.

cs.CV

Application of Machine Learning in understanding plant virus pathogenesis: Trends and perspectives on emergence, diagnosis, host-virus interplay and management

Inclusion of high throughput technologies in the field of biology has generated massive amounts of biological data in the recent years. Now, transforming these huge volumes of data into knowledge is the primary challenge in computational biology. The traditional methods of data analysis have failed to carry out the task. Hence, researchers are turning to machine learning based approaches for the analysis of high-dimensional big data. In machine learning, once a model is trained with a training dataset, it can be applied on a testing dataset which is independent. In current times, deep learning algorithms further promote the application of machine learning in several field of biology including plant virology. Considering a significant progress in the application of machine learning in understanding plant virology, this review highlights an introductory note on machine learning and comprehensively discusses the trends and prospects of machine learning in diagnosis of viral diseases, understanding host-virus interplay and emergence of plant viruses.

cs.LG

Nonlinear analysis of a classical double oscillator model

A classical double oscillator model, that includes in certain parameter limits, the standard harmonic oscillator and the inverse oscillator, is interpreted as a dynamical system. We study its essential features and make a qualitative analysis of orbits around the equilibrium points, period-doubling bifurcation, time series curves, surfaces of section and Poincare maps. An interesting outcome of our findings is the emergence of chaotic behavior when the system is confronted with a periodic force term like fcosωt.

physics.class-ph

Branched Hamiltonians for a quadratic type Liénard oscillator

We point out that when a quadratic type Liénard equation is suitably interpreted shows branching due to the double valuedness of the governing Hamiltonian. Under certain approximation of the guiding coupling constant we derive its quantum counterpart that is guided by a momentum-dependent mass function.

quant-ph

Classification of Cellular Automata Rules Based on Their Properties

This paper presents a classification of Cellular Automata rules based on its properties at the nth iteration. Elaborate computer program has been designed to get the nth iteration for arbitrary 1-D or 2-D CA rules. Studies indicate that the figures at some particular iteration might be helpful for some specific application. The hardware circuit implementation can be done using opto-electronic components [1-7].

cs.DM