arXiv Science⌕ Search

arXiv subjects

Md. Manzurul Hasan

Publications and source records attributed to Md. Manzurul Hasan.

10 recordsLinked to original sources

The Longest Common Bitonic Subsequence: Match-Sensitive Algorithms and Conditional Hardness

The longest common bitonic subsequence problem asks for a longest common subsequence of two ordered sequences whose values strictly increase and then strictly decrease; either phase may be empty. We formulate the problem through increasing and decreasing endpoint values at matching position pairs. This gives a constructive quadratic baseline and a matchsensitive algorithm based on two standard dominance-maximum passes. Its time is the sum of an input-sorting term and the number of matches times a squared logarithmic factor. We state the endpoint interface that permits reuse of increasing subsequence algorithms, and distinguish this specialization from new range searching machinery. A linear-size padding reduction transfers the conditional strongly subquadratic lower bound for longest common increasing subsequence to the bitonic problem. Reproducible implementations, exhaustive small-instance checks, and newly measured synthetic experiments document correctness and the practical tradeoff between sparse and dense processing.

cs.DS↗

ML-MAWS: Alignment-Free Maximum Likelihood Phylogeny Estimation Using Minimal Absent Words

Alignment-free methods in phylogenetic tree construction have major benefits in computational efficiency over alignment-based methods, but most sacrifice sequence information to pairwise distances, losing the statistical power of maximum likelihood (ML) inference. We describe ML-MAWS, an algorithm that fills this gap by encoding Minimal Absent Words (MAWs) as a binary presence/absence character matrix and estimating using an ML tree under the Lewis Mkv model using ascertainment bias correction. MAWs are obtained in linear time through the traversal of a suffix automaton. The pipeline incorporates strand-aware intersection filtering that retains only MAWs absent from both DNA orientations, entropy-based multi-length selection via Shannon entropy maximization to select the most informative lengths of MAWs, and parsimony-informative character capping to retain the most discriminative columns. We tested ML-MAWS on 14 benchmark datasets of bacterial, mitochondrial, viral, and simulated genomes with normalized Robinson-Foulds distances and matching split distances against published reference trees. The results show that while the binary encoding of MAWs can lead to higher topological error than continuous-valued distance baselines on closely related genomes, ML-MAWS is the first MAW-based method to provide per-branch bootstrap support and a rigorous probabilistic framework with ascertainment bias correction capabilities lacking from all existing alignment-free methods.

q-bio.PE↗

Novel hybrid protein scaffold gap filling using weighted machine learning ensemble, beam search, and mass-constrained reranking

Protein scaffold gap filling is an important computational task in protein sequence reconstruction, where missing amino acid regions must be inferred from incomplete scaffold information. This study proposes a hybrid machine learning and mass constrained reranking framework for protein scaffold gap filling under known-gap-size and known-gapmass settings. Homologous protein sequences from MabCampath, P5A proteoform, and carbonic anhydrase 2 were used to generate masked 11-mer residue-level samples and fullgap evaluation cases. The residue prediction task was formulated as a 20-class amino acid classification problem using first-, middle-, and last-position masking. Multiple classical machine learning models were trained using raw encoded, row-average, and SVD-reduced features, and the strongest models were combined through a validation-accuracy-weighted ensemble. For known-size gap reconstruction, beam search was used to generate complete missing peptide sequences from residue-level probability estimates. For known-mass reconstruction, mass-constrained homologous candidate retrieval was combined with hybrid reranking based on mass validity, homologous frequency, context support, ensemble likelihood, mass error, and length penalty. The proposed framework achieved 95.41% residue-level validation accuracy, 87.50% known-size exact-match accuracy, and 100% top-5 recovery on seven CAH2 known-mass benchmark cases. These results indicate that the proposed framework can effectively reconstruct missing protein regions by integrating local sequence learning, homologous evidence, peptide mass constraints, and biochemical validation.

q-bio.BM↗

G2P Explorer: A Native iOS Framework for Residue-Level Genomics to Proteomics Visualization and Structural Variant Interpretation

Genetic testing reports coding variants far faster than they can be interpreted, and placing a variant in its biophysical context, the domain it perturbs, whether its residue is buried or exposed, whether it lies near a disulfide bond or a predicted binding pocket increasingly requires projecting it onto a three dimensional protein model. The Genomics 2 Proteins (G2P) portal unifies the gene-to-structure identifier chain with a dense, residue indexed annotation table, but its visualization layer presumes a desktop browser and is awkward at the bedside, in the classroom, or in the field, where a phone or tablet is often the only device. We present G2P Explorer, a native iOS framework that consumes the public G2P REST API on device, parses its seventy one column tab separated feature tables without loss of fidelity, and presents the result through six interlinked modules sharing a single observable view model. Beyond porting, it contributes a SwiftUI Canvas multi-track sequence renderer, an on device reconstruction of the portal's unavailable isoform alignment route, a bidirectional Swift JavaScript structural bridge that absorbs the AlphaFold model file versioning scheme, and a fault tolerant ingestion layer that parses semi structured free text and distinguishes absent annotations from zero. Each searched protein is cached on device after the first fetch, so it reopens instantly and works offline, and the embedded structural view is drawn in a reduced form suited to a small screen. Across six proteins spanning 189-1{,}863 residues, the framework sustains interactive frame times (3.1-16.6\,ms) and modest memory (16-72\,MB). G2P Explorer is an open, reproducible mobile companion to the G2P portal for hypothesis generation, teaching, and on the go variant interpretation.

q-bio.QM↗

Expected Cost of Greedy Online Facility Assignment on Regular Polygons (v3)

We study a greedy online facility assignment process on a regular $n$-gon, where unit-capacity facilities occupy the vertices and customers arrive sequentially at uniformly random locations on polygon edges. Each arrival is irrevocably assigned to the nearest currently free facility under the shortest edge-walk metric, with uniform tie-breaking among equidistant choices. Our main theoretical result is an exact value-function characterization: for every occupancy state $S\subseteq V$, the expected remaining cost $V(S)$ satisfies a finite-horizon integral recurrence obtained by conditioning on the random arrival edge and position. To make this recurrence computationally effective, we exploit dihedral symmetry of the regular polygon and show that $V(S)$ is invariant under rotations and reflections, enabling canonicalization and symmetry-reduced dynamic programming. For small $n$, we evaluate the recurrence accurately using deterministic numerical integration over piecewise-linear distance regions,; for larger $n$, we estimate the expected total cost via direct Monte Carlo simulation of the online process and report $95\%$ confidence intervals. Our computations validate the recurrence (including a closed-form check for the square, $n=4$) and indicate that the total expected cost increases with $n$, while the per-customer expected travel distance grows gradually as remaining free vertices become farther on average. \keywords{Online algorithms \and Facility assignment \and Expected cost \and Regular polygons \and Symmetry reduction \and Monte Carlo}

cs.DS↗

Permutation Matching Under Parikh Budgets: Linear-Time Detection, Packing, and Disjoint Selection

We study permutation (jumbled/Abelian) pattern matching over a general alphabet $Σ$. Given a pattern P of length m and a text T of length n, the classical task is to decide whether T contains a length-m substring whose Parikh vector equals that of P . While this existence problem admits a linear-time sliding-window solution, many practical applications require optimization and packing variants beyond mere detection. We present a unified sliding-window framework based on maintaining the Parikh-vector difference between P and the current window of T , enabling permutation matching in O(n + σ) time and O(σ) space, where σ = |Σ|. Building on this foundation, we introduce a combinatorial-optimization variant that we call Maximum Feasible Substring under Pattern Supply (MFSP): find the longest substring S of T whose symbol counts are component-wise bounded by those of P . We show that MFSP can also be solved in O(n + σ) time via a two-pointer feasibility maintenance algorithm, providing an exact packing interpretation of P as a resource budget. Finally, we address non-overlapping occurrence selection by modeling each permutation match as an equal-length interval and proving that a greedy earliest-finishing strategy yields a maximum-cardinality set of disjoint matches, computable in linear time once all matches are enumerated. Our results provide concise, provably correct algorithms with tight bounds, and connect frequency-based string matching to packing-style optimization primitives.

cs.DS↗

OBHS: An Optimized Block Huffman Scheme for Real-Time Audio Compression

In this paper, we introduce OBHS (Optimized Block Huffman Scheme), a novel lossless audio compression algorithm tailored for real-time streaming applications. OBHS leverages block-wise Huffman coding with canonical code representation and intelligent fallback mechanisms to achieve high compression ratios while maintaining low computational complexity. Our algorithm partitions audio data into fixed-size blocks, constructs optimal Huffman trees for each block, and employs canonical codes for efficient storage and transmission. Experimental results demonstrate that OBHS attains compression ratios of up to 93.6% for silence-rich audio and maintains competitive performance across various audio types, including pink noise, tones, and real-world recordings. With a linear time complexity of O(n) for n audio samples, OBHS effectively balances compression efficiency and computational demands, making it highly suitable for resource-constrained real-time audio streaming scenarios.

cs.SD↗

A Space-Efficient Algorithm for Longest Common Almost Increasing Subsequence of Two Sequences

Let $A$ and $B$ be two number sequences of length $n$ and $m$, respectively, where $m\le n$. Given a positive number $δ$, a common almost increasing sequence $s_1\ldots s_k$ is a common subsequence for both $A$ and $B$ such that for all $2\le i\le k$, $s_i+δ> \max_{1\le j < i} s_j$. The LCaIS problem seeks to find the longest common almost increasing subsequence (LCaIS) of $A$ and $B$. An LCaIS can be computed in $O(nm\ell)$ time and $O(nm)$ space [Ta, Shieh, Lu (TCS 2021)], where $\ell$ is the length of the LCaIS of $A$ and $B$. In this paper we first give an $O(nm\ell)$-time and $O(n+m\ell)$-space algorithm to find LCaIS, which improves the space complexity. We then design an $O((n+m)\log n +\mathcal{M}\log \mathcal{M} + \mathcal{C}\ell)$-time and $O(\mathcal{M}(\ell+\log \mathcal{M}))$-space algorithm, which is faster when the number of matching pairs $\mathcal{M}$ and the number of compatible matching pairs $\mathcal{C}$ are in $o(nm/\log m)$.

cs.DS↗

Online Facility Assignments on Polygons

We study the online facility assignment problem on regular polygons, where all sides are of equal length. The influence of specific geometric settings has remained mostly unexplored, even though classical online facility assignment problems have mainly dealt with linear and general metric spaces. We fill this gap by considering the following four basic geometric settings: equilateral triangles, rectangles, regular $n$-polygons, and circles. The facilities are situated at fixed positions on the boundary, and customers appear sequentially on the boundary. A customer needs to be assigned immediately without any information about future customer arrivals. We study a natural greedy algorithm. First, we study an equilateral triangle with three facilities at its corners; customers can appear anywhere on the boundary. We then analyze regular $n$-sided polygons, obtaining a competitive ratio of $2n-1$, showing that the algorithm performance degrades linearly with the number of corner points for polygons. For the circular configuration, the competitive ratio is $2n-1$ when the distance between two adjacent facilities is the same. And the competitive ratios are $n^2-n+1$ and $2^n - 1$ for varying distances linearly and exponentially respectively. Each facility has a fixed capacity proportional to the geometric configuration, and customers appear only along the boundary edges. Our results also show that simpler geometric configurations have more efficient performance bounds and that spacing facilities uniformly apart prevent worst-case scenarios. The findings have many practical implications because large networks of facilities are best partitioned into smaller and geometrically simple pieces to guarantee good overall performance.

cs.DS↗

Positive Planar Satisfiability Problems under 3-Connectivity Constraints

A 3-SAT problem is called positive and planar if all the literals are positive and the clause-variable incidence graph (i.e., SAT graph) is planar. The NAE 3-SAT and 1-in-3-SAT are two variants of 3-SAT that remain NP-complete even when they are positive. The positive 1-in-3-SAT problem remains NP-complete under planarity constraint, but planar NAE 3-SAT is solvable in $O(n^{1.5}\log n)$ time. In this paper we prove that a positive planar NAE 3-SAT is always satisfiable when the underlying SAT graph is 3-connected, and a satisfiable assignment can be obtained in linear time. We also show that without 3-connectivity constraint, existence of a linear-time algorithm for positive planar NAE 3-SAT problem is unlikely as it would imply a linear-time algorithm for finding a spanning 2-matching in a planar subcubic graph. We then prove that positive planar 1-in-3-SAT remains NP-complete under the 3-connectivity constraint, even when each variable appears in at most 4 clauses. However, we show that the 3-connected planar 1-in-3-SAT is always satisfiable when each variable appears in an even number of clauses.

cs.CC↗