arXiv ScienceSearch

arXiv subjects

Shilun Li

Publications and source records attributed to Shilun Li.

12 recordsLinked to original sources

A Public-Key-Dependent Adversarial-Deletion Ceiling for Fixed-Alphabet Multi-Bit Pseudorandom Codes

A pseudorandom code (PRC) is a keyed error-correcting code whose codewords are computationally indistinguishable from uniform strings. We study public-key PRCs over fixed alphabets against adversarial deletions, where the deletion channel may both depend on the public encoding key and the transmitted codeword. Let $γ_q^{\mathrm{LCS}}$ denote the asymptotic normalised longest-common-subsequence length of two independent uniform $q$-ary strings. We prove that for every fixed $q\ge2$ no multi-message public-key PRC with a single-output decoder is robust against all such $δ$-deletion channels for any $δ>1-γ_q^{\mathrm{LCS}}$. For $q=2$, the current rigorous bound $γ_2\ge0.792665992$ rules out every constant $δ>0.207334008$. The proof uses pseudorandomness only to transfer an LCS event from uniform strings to independently sampled codewords and therefore also the resulting collision argument is information-theoretic and requires no secret key. We also extend the argument to list decoding: for every fixed constant $L$, provided the message space contains at least $L+1$ messages, no such PRC with output lists of size at most $L$ is robust for $δ>1-γ_2^{(L+1)}$. Since $γ_2^{(m)}=1/2+Θ(1/\sqrt m)$, these thresholds approach $1/2$. Our bounds are specific to public-key-dependent adversarial channels and do not apply to oblivious edit channels.

cs.CR

Average-Radius List-Decodability of Random Linear Codes

We prove that for every prime power $q$ and every $p \in (0, 1-1/q)$, a random $\mathbb{F}_q$-linear code of rate $1 - h_q(p) - ε$ is $(p, C_{p,q}/ε)$-average-radius list-decodable with probability at least $1 - q^{-Ω(n)}$, i.e., for every center $y \in \mathbb{F}_q^n$, the $C_{p,q}/ε$ codewords closest to $y$ have average fractional Hamming distance at least $p$ from $y$. This extends a similar result for (standard) list-decoding due to Guruswami, Håstad, and Kopparty (2010) to the stronger average-radius guarantee, with the same $O(1/ε)$ list size. For average-radius list-decoding, such a result was previously known only for binary linear codes (Guruswami, Li, Mosheiff, Resch, Silas, and Wootters, 2021) and for general (non-linear) random codes over arbitrary alphabets (Elias, 1991).

cs.IT

A Symplectic Proof of the Quantum Singleton Bound

We present a symplectic linear-algebraic proof of the Quantum Singleton Bound for stabiliser quantum error-correcting codes together with a Lean4 formalisation of the linear-algebraic argument. The proof is formulated in the language of finite-dimensional symplectic vector spaces modelling Pauli operators and relies on distance-based erasure correctability and the cleaning lemma. Using a dimension-counting argument within the symplectic stabiliser framework, we derive the bound $k + 2(d-1) \le n$ for any $[[n, k, d]]$ stabiliser code. This approach isolates the algebraic structure underlying the bound and avoids the heavier analytic machinery that appears in entropy-based proofs, while remaining well-suited to formal verification.

quant-ph

Density Frankl-Rödl on the Sphere

We establish a density variant of the Frankl-Rödl theorem on the sphere $\mathbb{S}^{n-1}$, which concerns avoiding pairs of vectors with a specific distance, or equivalently, a prescribed inner product. In particular, we establish lower bounds on the probability that a randomly chosen pair of such vectors lies entirely within a measurable subset $A \subseteq \mathbb{S}^{n-1}$ of sufficiently large measure. Additionally, we prove a density version of spherical avoidance problems, which generalize from pairwise avoidance to broader configurations with prescribed pairwise inner products. Our framework encompasses a class of configurations we call inductive configurations, which include simplices with any prescribed inner product $-1 < r < 1$. As a consequence of our density statement, we show that all inductive configurations are sphere Ramsey.

math.PR

A Deterministic Construction of a Large Distance Code from the Wozencraft Ensemble

We present an explicit construction of a sequence of rate $1/2$ Wozencraft ensemble codes (over any fixed finite field $\mathbb{F}_q$) that achieve minimum distance $Ω(\sqrt{k})$ where $k$ is the message length. The coefficients of the Wozencraft ensemble codes are constructed using Sidon Sets and the cyclic structure of $\mathbb{F}_{q^{k}}$ where $k+1$ is prime with $q$ a primitive root modulo $k+1$. Assuming Artin's conjecture, there are infinitely many such $k$ for any prime power $q$.

cs.IT

Dynamics and Probability in the Toss of a Coin with Symmetric Inhomogeneous Density

Under investigation in this paper is the dynamics and probability of heads in the toss of a coin with symmetric inhomogeneous density. Such coins are assumed to have diagonal inertia matrix. The rotational motion of the coin is determined by the initial angular momentum and initial position of the coin. We described the dynamic behavior of the unit normal vector and calculated the limiting probability of heads as time goes to infinity with respect to the fixed initial parameters. Our probability formula extends the formula for homogeneous coins by Keller and Diaconis et al.

math.PR

On the Problem of Undirected st-connectivity

In this paper, we discuss an algorithm for the problem of undirected st-connectivity that is deterministic and log-space, namely that of Reingold within his 2008 paper "Undirected Connectivity in Log-Space". We further present a separate proof by Rozenman and Vadhan of $\text{USTCONN} \in L$ and discuss its similarity with Reingold's proof. Undirected st-connectively is known to be complete for the complexity class SL--problems solvable by symmetric, non-deterministic, log-space algorithms. Likewise, by Aleliunas et. al., it is known that undirected st-connectivity is within the RL complexity class, problems solvable by randomized (probabilistic) Turing machines with one-sided error in logarithmic space and polynomial time. Finally, our paper also shows that undirected st-connectivity is within the L complexity class, problems solvable by deterministic Turing machines in logarithmic space. Leading from this result, we shall explain why SL = L and discuss why is it believed that RL = L.

cs.CC

Indistinguishability Obfuscation of Circuits and its Application in Security

Under discussion in the paper is an $i\mathcal{O}$ (indistinguishability obfuscator) for circuits in Nick's Class. The obfuscator is constructed by encoding the Branching Program given by Barrington's theorem using Multilinear Jigsaw Puzzle framework. We will show under various indistinguishability hardness assumptions, the constructed obfuscator is an $i\mathcal{O}$ for Nick's Class. Using Fully Homomorphic Encryption, we will amplify the result and construct an $i\mathcal{O}$ for $\textbf{P}/poly$, which are circuits of polynomial size. Discussion on $i\mathcal{O}$ and Functional Encryption is also included in this paper.

cs.CR

Distributionally Robust Classifiers in Sentiment Analysis

In this paper, we propose sentiment classification models based on BERT integrated with DRO (Distributionally Robust Classifiers) to improve model performance on datasets with distributional shifts. We added 2-Layer Bi-LSTM, projection layer (onto simplex or Lp ball), and linear layer on top of BERT to achieve distributionally robustness. We considered one form of distributional shift (from IMDb dataset to Rotten Tomatoes dataset). We have confirmed through experiments that our DRO model does improve performance on our test set with distributional shift from the training set.

cs.CL

Playing 2048 With Reinforcement Learning

The game of 2048 is a highly addictive game. It is easy to learn the game, but hard to master as the created game revealed that only about 1% games out of hundreds million ever played have been won. In this paper, we would like to explore reinforcement learning techniques to win 2048. The approaches we have took include deep Q-learning and beam search, with beam search reaching 2048 28.5 of time.

cs.AI

Ensemble ALBERT on SQuAD 2.0

Machine question answering is an essential yet challenging task in natural language processing. Recently, Pre-trained Contextual Embeddings (PCE) models like Bidirectional Encoder Representations from Transformers (BERT) and A Lite BERT (ALBERT) have attracted lots of attention due to their great performance in a wide range of NLP tasks. In our Paper, we utilized the fine-tuned ALBERT models and implemented combinations of additional layers (e.g. attention layer, RNN layer) on top of them to improve model performance on Stanford Question Answering Dataset (SQuAD 2.0). We implemented four different models with different layers on top of ALBERT-base model, and two other models based on ALBERT-xlarge and ALBERT-xxlarge. We compared their performance to our baseline model ALBERT-base-v2 + ALBERT-SQuAD-out with details. Our best-performing individual model is ALBERT-xxlarge + ALBERT-SQuAD-out, which achieved an F1 score of 88.435 on the dev set. Furthermore, we have implemented three different ensemble algorithms to boost overall performance. By passing in several best-performing models' results into our weighted voting ensemble algorithm, our final result ranks first on the Stanford CS224N Test PCE SQuAD Leaderboard with F1 = 90.123.

cs.CL

Trajectory Prediction using Generative Adversarial Network in Multi-Class Scenarios

Predicting traffic agents' trajectories is an important task for auto-piloting. Most previous work on trajectory prediction only considers a single class of road agents. We use a sequence-to-sequence model to predict future paths from observed paths and we incorporate class information into the model by concatenating extracted label representations with traditional location inputs. We experiment with both LSTM and transformer encoders and we use generative adversarial network as introduced in Social GAN to learn the multi-modal behavior of traffic agents. We train our model on Stanford Drone dataset which includes 6 classes of road agents and evaluate the impact of different model components on the prediction performance in multi-class scenes.

cs.LG