arXiv ScienceSearch

arXiv subjects

James Gleeson

Publications and source records attributed to James Gleeson.

7 recordsLinked to original sources

Landau theory applied to antiferroelectric ordering in ferroelectric nematic liquid crystals

The polarization and density modulation associated with antiferroelectric ordering is studied experimentally as a function of temperature in two ferroelectric nematic liquid crystals, the prototypical single compound (DIO) and a commercial mixture (FNLC919). The modulation wavenumber qA is determined by small angle X-ray diffraction from the weak smectic-like density wave (wavenumber qS = 2qA) that accompanies the polarization modulation. Results for qS and the saturated value of the polarization are analyzed in terms of Landau theory previously developed to describe the para-/antiferro-/feroelectric sequence of phase transitions in solid ferroelectrics. The analysis indicates that the polarization modulation is reasonably well approximated by a simple sinusoid in the antiferroelectric phase of DIO, whereas in FNLC919 the modulation develops a strongly soliton-like profile (with sharply decreasing wavenumber) close to the antiferro- to ferrolectric transition.

cond-mat.soft

Director-layer dynamics in the antiferroelectric smectic-ZA phase of a ferroelectric nematic liquid crystal

A dynamic light scattering study of director-layer fluctuations in the antiferroelectric smectic-ZA phase of the ferroelectric nematic liquid crystal DIO is reported. The dynamics are consistent with the distinctive feature of the ZA phase that the smectic layers form parallel to the axis of molecular orientational order (director). A model is developed to describe quantitatively the dispersion of the fluctuation relaxation rates. The model is based on a specialization of the elastic free energy density of the smectic-C phase to the case of 90 degree director tilt, a "first-order" approximation of the viscous stresses by their form for an incompressible uniaxial fluid, and a treatment of the effect of chevron layer structure that develops in planar sample cells due to temperature-dependent layer shrinkage, as documented in previous studies on DIO. From the modeling, the layer compression elastic constant is estimated to be ~100 times lower in the smectic-ZA phase than in an ordinary smectic-A liquid crystal. Possible effects of the antiferroelectric layer polarization on the director splay elasticity and viscosity are described. The temperature dependencies of the splay, twist, and bend elastic constants and associated viscosities in the higher temperature nematic phase are also presented.

cond-mat.soft

Minuet: Accelerating 3D Sparse Convolutions on GPUs

Sparse Convolution (SC) is widely used for processing 3D point clouds that are inherently sparse. Different from dense convolution, SC preserves the sparsity of the input point cloud by only allowing outputs to specific locations. To efficiently compute SC, prior SC engines first use hash tables to build a kernel map that stores the necessary General Matrix Multiplication (GEMM) operations to be executed (Map step), and then use a Gather-GEMM-Scatter process to execute these GEMM operations (GMaS step). In this work, we analyze the shortcomings of prior state-of-the-art SC engines, and propose Minuet, a novel memory-efficient SC engine tailored for modern GPUs. Minuet proposes to (i) replace the hash tables used in the Map step with a novel segmented sorting double-traversed binary search algorithm that highly utilizes the on-chip memory hierarchy of GPUs, (ii) use a lightweight scheme to autotune the tile size in the Gather and Scatter operations of the GMaS step, such that to adapt the execution to the particular characteristics of each SC layer, dataset, and GPU architecture, and (iii) employ a padding-efficient GEMM grouping approach that reduces both memory padding and kernel launching overheads. Our evaluations show that Minuet significantly outperforms prior SC engines by on average $1.74\times$ (up to $2.22\times$) for end-to-end point cloud network executions. Our novel segmented sorting double-traversed binary search algorithm achieves superior speedups by $15.8\times$ on average (up to $26.8\times$) over prior SC engines in the Map step. The source code of Minuet is publicly available at https://github.com/UofT-EcoSystem/Minuet.

cs.DC

Optimizing Data Collection in Deep Reinforcement Learning

Reinforcement learning (RL) workloads take a notoriously long time to train due to the large number of samples collected at run-time from simulators. Unfortunately, cluster scale-up approaches remain expensive, and commonly used CPU implementations of simulators induce high overhead when switching back and forth between GPU computations. We explore two optimizations that increase RL data collection efficiency by increasing GPU utilization: (1) GPU vectorization: parallelizing simulation on the GPU for increased hardware parallelism, and (2) simulator kernel fusion: fusing multiple simulation steps to run in a single GPU kernel launch to reduce global memory bandwidth requirements. We find that GPU vectorization can achieve up to $1024\times$ speedup over commonly used CPU simulators. We profile the performance of different implementations and show that for a simple simulator, ML compiler implementations (XLA) of GPU vectorization outperform a DNN framework (PyTorch) by $13.4\times$ by reducing CPU overhead from repeated Python to DL backend API calls. We show that simulator kernel fusion speedups with a simple simulator are $11.3\times$ and increase by up to $1024\times$ as simulator complexity increases in terms of memory bandwidth requirements. We show that the speedups from simulator kernel fusion are orthogonal and combinable with GPU vectorization, leading to a multiplicative speedup.

cs.LG

RL-Scope: Cross-Stack Profiling for Deep Reinforcement Learning Workloads

Deep reinforcement learning (RL) has made groundbreaking advancements in robotics, data center management and other applications. Unfortunately, system-level bottlenecks in RL workloads are poorly understood; we observe fundamental structural differences in RL workloads that make them inherently less GPU-bound than supervised learning (SL). To explain where training time is spent in RL workloads, we propose RL-Scope, a cross-stack profiler that scopes low-level CPU/GPU resource usage to high-level algorithmic operations, and provides accurate insights by correcting for profiling overhead. Using RL-Scope, we survey RL workloads across its major dimensions including ML backend, RL algorithm, and simulator. For ML backends, we explain a $2.3\times$ difference in runtime between equivalent PyTorch and TensorFlow algorithm implementations, and identify a bottleneck rooted in overly abstracted algorithm implementations. For RL algorithms and simulators, we show that on-policy algorithms are at least $3.5\times$ more simulation-bound than off-policy algorithms. Finally, we profile a scale-up workload and demonstrate that GPU utilization metrics reported by commonly used tools dramatically inflate GPU usage, whereas RL-Scope reports true GPU-bound time. RL-Scope is an open-source tool available at https://github.com/UofT-EcoSystem/rlscope .

cs.LG

The interplay between spatial and heliconical bond order in twist-bend nematic materials

The nanostructure of two novel sulfur containing dimer materials has been investigated experimentally by hard and by resonant tender X-ray scattering techniques. On cooling the dimers through the nematic to twist-bend nematic (N-NTB) phase transition, the correlation length associated with short-range positional order drops, while the heliconical orientational order becomes more correlated. The heliconical pitch shows a stronger temperature dependence near the N-NTB transition than observed in previously studied dimers, such as the CBnCB series of compounds. We explain both this strong variation and the dependence of the heliconical pitch on the length of the spacer connecting the monomer units by taking into account a temperature dependent molecular bend and intermolecular overlap. and. The heliconical structure is observed even in the upper 3-4{\deg}C range of the smectic phase that forms just below the NTB state. The coexistence of smectic layering and the heliconical order indicates a SmCTB -type phase where the rigid units of the dimers are tilted with respect to the layer normal in order to accommodate the bent conformation of the dimers, but the tilt direction rotates along the heliconical axis. This is potentially similar to the SmCTB phase reported by Abberley et al (Nat. Commun. 2018, 9, 228) below a SmA phase.

cond-mat.soft

Prediction with Restricted Resources and Finite Automata

We obtain an index of the complexity of a random sequence by allowing the role of the measure in classical probability theory to be played by a function we call the generating mechanism. Typically, this generating mechanism will be a finite automata. We generate a set of biased sequences by applying a finite state automata with a specified number, $m$, of states to the set of all binary sequences. Thus we can index the complexity of our random sequence by the number of states of the automata. We detail optimal algorithms to predict sequences generated in this way.

stat.ML