arXiv ScienceSearch

arXiv subjects

Mingwen Zhang

Publications and source records attributed to Mingwen Zhang.

9 recordsLinked to original sources

Multi-Agent Self-Improving Reinforcement Learning for Video Reasoning

Video reasoning tasks such as grounded video question answering and temporal grounding require selecting temporal evidence that supports the query. In many current training setups, temporal supervision is applied through local objectives such as boundary regression or span generation, while verification is used mainly to rerank candidate segments at inference time. We study whether a frozen verifier can also guide training. Our multi-agent framework couples a trainable \emph{Grounder} with a frozen \emph{Verifier}: the Grounder samples candidate trajectories and evidence segments, the Verifier assigns query-conditioned segment scores, a group-relative policy-gradient objective favors trajectories that outperform their within-input peers, and a bootstrapped calibration loss steers temporal predictions toward verifier-preferred spans. Trained on source tasks and evaluated without target-dataset fine-tuning, a two-billion-parameter instantiation transfers zero-shot across grounded question answering, temporal grounding, and long-video question answering, reaching 28.7\% intersection-over-union and 25.4\% answer-grounding accuracy on a grounded-question-answering benchmark, 46.1\% intersection-over-union on a temporal-grounding benchmark, and 54.1\% on a long-video question-answering benchmark. Relative to a strong same-scale baseline, the gains are modest but consistent, with the clearest improvements on relevance-oriented metrics such as intersection-over-union and moderate-overlap recall. Within the tested benchmarks and transfer setting, the results support frozen verification as a training signal for evidence selection, while showing that strict boundary precision remains comparatively weaker. Code and models are available at https://anonymous.4open.science/r/MASIRL-E50C/

cs.CV

The Art of Socratic Inquiry: A Framework for Proactive Template-Guided Therapeutic Conversation Generation

Proactive questioning, where therapists deliberately initiate structured, cognition-guiding inquiries, is a cornerstone of cognitive behavioral therapy (CBT). Yet, current psychological large language models (LLMs) remain overwhelmingly reactive, defaulting to empathetic but superficial responses that fail to surface latent beliefs or guide behavioral change. To bridge this gap, we propose the \textbf{Socratic Inquiry Framework (SIF)}, a lightweight, plug-and-play therapeutic intent planner that transforms LLMs from passive listeners into active cognitive guides. SIF decouples \textbf{when to ask} (via Strategy Anchoring) from \textbf{what to ask} (via Template Retrieval), enabling context-aware, theory-grounded questioning without end-to-end retraining. Complementing SIF, we introduce \textbf{Socratic-QA}, a high-quality dataset of strategy-aligned Socratic sequences that provides explicit supervision for proactive reasoning. Experiments show that SIF significantly enhances proactive questioning frequency, conversational depth, and therapeutic alignment, marking a clear shift from reactive comfort to proactive exploration. Our work establishes a new paradigm for psychologically informed LLMs: not just to respond, but to guide.

cs.CL

On-chip quadratically nonlinear photodetector

Involving deterministically nonlinear photoresponse in on-chip photodetector is intriguing to develop sophisticated functions in photonic integrated circuits, such as in-sensor computing and optoelectronic mixing, though the corresponding devices are still lack of sufficient investigation. Here, we demonstrate an on-chip quadratically nonlinear photodetector (QNPD) by configuring an InSe p-i-n homojunction on a silicon waveguide. Telecom-band light guiding in the waveguide couples with the InSe evanescently and is frequency up-converted into visible light via InSe's second-harmonic generation (SHG), which is subsequently absorbed by InSe and finally generates photocurrent under the built-in electric field of the p-i-n homojunction. Governed by these sequential processes, the on-chip QNPD presents a quadratic function between photocurrent and optical power. Thanks to the efficient SHG and well-established homojunction in InSe, the QNPD reaches a high normalized responsivity of 37.1 A/W2 and low dark current of 1 pA, representing greatly improved performances among reported nonlinear photodetectors. Benefiting from the extra SHG process, the on-chip QNPD intrinsically incorporates light-light interactions, enabling straightforwardly monitoring all-optically mixing signals electrically. As an example, an array of 16-pixel QNPDs was designed to implement a fully single-shot on-chip autocorrelator without requirement of bulky optics and external cameras, which precisely measures picosecond pulses with high sensitivity of 6.1*10-10 W2.

physics.optics

Nonlinear photodetector based on InSe p-n homojunction for improving spatial imaging resolution

We demonstrate an efficient nonlinear photodetector (NLPD) with quadratic response based on a few-layer InSe p-n homojunction, which is beneficial from the strong second harmonic generation (SHG) process in InSe and effective harvest of photocarriers actuated by the high-quality homojunction. The NLPD can sense light with photon energy smaller than InSe electronic bandgap because the SHG process in InSe doubles the frequency of incident light, extending InSe photodetection wavelength range to 1750 nm. The InSe p-n homojunction, which is electrostatically doped by two split back gates, presents a rectification ratio exceeding 106 with a dark current down to 2 pA and a high normalized responsivity of 0.534 A/W2 for the telecom-band pulsed light at 1550 nm. The photocurrents of the SHG-assisted photodetection have a quadratic dependence on the optical powers, making the NLPD highly sensitive to light intensity variation with improved spatial resolution. As examples, the NLPD is employed to precisely determine the localization point of a focused laser beam waist and implement spatial imaging with an improved resolution compared with the linear photodetector. These features highlight the potential of the proposed NLPD in developing advanced optical sensing and imaging systems.

physics.optics

Approaching the robust linearity in dual-floating van der Waals photodiode

Two-dimensional (2D) material photodetectors have gained great attention as potential elements for optoelectronic applications. However, the linearity of the photoresponse is often compromised by the carrier interaction, even in 2D photodiodes. In this study, we present a new device concept of dual-floating van der Waals heterostructures (vdWHs) photodiode by employing ambipolar MoTe2 and n-type MoS2 2D semiconductors. The presence of type II heterojunctions on both sides of channel layers effectively deplete carriers and restrict the photocarrier trapping within the channel layers. As a result, the device exhibits robust linear photoresponse under photovoltaic mode from the visible (405 nm) to near-infrared (1600 nm) band. With the built-in electric field of the vdWHs, we achieve a linear dynamic range of ~ 100 dB, responsivity of ~ 1.57 A/W, detectivity of ~ 4.28 * 10^11 Jones, and response speed of ~ 30 {\mu}s. Our results showcase a promising device concept with excellent linearity towards fast and low-loss detection, high-resolution imaging, and logic optoelectronics.

cond-mat.mes-hall

Strong Second Harmonic Generation from Bilayer Graphene with Symmetry Breaking by Redox-Governed Charge Doping

Missing second-order nonlinearity in centrosymmetric graphene overshadows its intriguing optical attribute. Here, we report redox-governed charge doping could effectively break the centrosymmetry of bilayer graphene (BLG), enabling a strong second harmonic generation (SHG) with a strength close to that of the well-known monolayer MoS2. Verified from control experiments with in situ electrical current annealing and electrically gate-controlled SHG, the required centrosymmetry breaking of the emerging SHG arises from the charge-doping on the bottom layer of BLG by the oxygen/water redox couple. Our results not only reveal that charge doping is an effective way to break the inversion symmetry of BLG despite its strong interlayer coupling but also indicate that SHG spectroscopy is a valid technique to probe molecular doping on two-dimensional materials.

physics.optics

Electrically tunable second harmonic generation in atomically thin ReS2

Electrical tuning of second-order nonlinearity in optical materials is attractive to strengthen and expand the functionalities of nonlinear optical technologies, though its implementation remains elusive. Here, we report the electrically tunable second-order nonlinearity in atomically thin ReS2 flakes benefiting from their distorted 1T crystal structure and interlayer charge transfer. Enabled by the efficient electrostatic control of the few-atomic-layer ReS2, we show that second harmonic generation (SHG) can be induced in odd-number-layered ReS2 flakes which are centrosymmetric and thus without intrinsic SHG. Moreover, the SHG can be precisely modulated by the electric field, reversibly switching from almost zero to an amplitude more than one order of magnitude stronger than that of the monolayer MoS2. For the even-number-layered ReS2 flakes with the intrinsic SHG, the external electric field could be leveraged to enhance the SHG. We further perform the first-principles calculations which suggest that the modification of in-plane second-order hyperpolarizability by the redistributed interlayer-transferring charges in the distorted 1T crystal structure underlies the electrically tunable SHG in ReS2. With its active SHG tunability while using the facile electrostatic control, our work may further expand the nonlinear optoelectronic functions of two-dimensional materials for developing electrically controllable nonlinear optoelectronic devices.

physics.optics

ForgeryNet -- Face Forgery Analysis Challenge 2021: Methods and Results

The rapid progress of photorealistic synthesis techniques has reached a critical point where the boundary between real and manipulated images starts to blur. Recently, a mega-scale deep face forgery dataset, ForgeryNet which comprised of 2.9 million images and 221,247 videos has been released. It is by far the largest publicly available in terms of data-scale, manipulations (7 image-level approaches, 8 video-level approaches), perturbations (36 independent and more mixed perturbations), and annotations (6.3 million classification labels, 2.9 million manipulated area annotations, and 221,247 temporal forgery segment labels). This paper reports methods and results in the ForgeryNet - Face Forgery Analysis Challenge 2021, which employs the ForgeryNet benchmark. The model evaluation is conducted offline on the private test set. A total of 186 participants registered for the competition, and 11 teams made valid submissions. We will analyze the top-ranked solutions and present some discussion on future work directions.

cs.CV

Tips and Tricks for Webly-Supervised Fine-Grained Recognition: Learning from the WebFG 2020 Challenge

WebFG 2020 is an international challenge hosted by Nanjing University of Science and Technology, University of Edinburgh, Nanjing University, The University of Adelaide, Waseda University, etc. This challenge mainly pays attention to the webly-supervised fine-grained recognition problem. In the literature, existing deep learning methods highly rely on large-scale and high-quality labeled training data, which poses a limitation to their practicability and scalability in real world applications. In particular, for fine-grained recognition, a visual task that requires professional knowledge for labeling, the cost of acquiring labeled training data is quite high. It causes extreme difficulties to obtain a large amount of high-quality training data. Therefore, utilizing free web data to train fine-grained recognition models has attracted increasing attentions from researchers in the fine-grained community. This challenge expects participants to develop webly-supervised fine-grained recognition methods, which leverages web images in training fine-grained recognition models to ease the extreme dependence of deep learning methods on large-scale manually labeled datasets and to enhance their practicability and scalability. In this technical report, we have pulled together the top WebFG 2020 solutions of total 54 competing teams, and discuss what methods worked best across the set of winning teams, and what surprisingly did not help.

cs.CV