arXiv Science⌕ Search

arXiv subjects

Vladimir Beskorovainyi

Publications and source records attributed to Vladimir Beskorovainyi.

4 recordsLinked to original sources

Machine Learning for Coding Retail Product Names to Consumer-Price Categories: A Rule-plus-Bag-of-Words Pipeline with Reliability-Weighted Human-in-the-Loop Labeling

Price statistics increasingly draw on scanner, web-scraped and receipt data, whose product descriptions are short, noisy and carry no standard product code, so each item must be coded to a consumption classification such as COICOP. National statistical offices already report that lightweight text classifiers are adequate for this task, but the published evidence is thin on dispersion, paired comparison, train-test overlap and measured computational cost. This paper supplies that evaluation. On an openly released synthetic benchmark of six COICOP-like categories, seven models are trained under one matched protocol and compared on accuracy and cost. A character n-gram logistic regression is the most accurate model in every category (mean F1 = 0.997, and 0.996 on test items unseen in training). The small CNN and LSTM examined here never win a comparison, but once trained to the same stopping criterion they are not reliably worse than the word-level models either; what separates them is cost, at 65 and 77 times the training time of the cheapest model. The most accurate model is not the cheapest: character features cut inference throughput to 44,900 items per second against 212,700 for unigram bag-of-words. A rule-based prefix-tree stage admits 75-86% of positive items, so it bounds the cascade's recall. A Monte Carlo study of the labeling protocol, in which annotators are simulated and no human annotation was collected, shows that an additive reliability weight barely improves on majority vote while Dawid-Skene aggregation recovers labels markedly better, and that the difference carries through to classifiers trained on those labels. The benchmark is synthetic because the production data behind the architecture are confidential; its generator applies character-level corruption and shares phrase sets with the rule stage, and both conditions are stated wherever a result depends on them.

cs.CL↗

A machine-readable catalogue of the Tsiolkovsky papers (fond 555, Archive of the Russian Academy of Sciences), and a way to measure how well its handwriting can be read

The personal archive of Konstantin Tsiolkovsky (1857-1935) is held as fond 555 of the Archive of the Russian Academy of Sciences. The archive scanned the fond and published the images, but with no queryable catalogue, no full-text search and no dataset: the holdings can only be browsed one page at a time. This paper describes a machine-readable catalogue of all 2,019 files and 51,008 scans, a dating for 1,969 files taken from the archive's own descriptions, a page-level classification of every scan into handwriting and typescript, and a machine transcription of the fond in full: all 2,019 files and all 51,008 scans. It also reports a way to measure handwritten-text-recognition accuracy in an archive with no ground truth. Archives of the typewriter era often preserve one text twice, as manuscript and as a typed copy; transcribing both and comparing isolates the reading error, since source and pipeline are identical and only page difficulty differs. Across 1,759 such pairs from 224 files, two readings of a handwritten page agree on a median 37% of words and share a longest verbatim run of a median 10 words; the median is unchanged from the 294 pairs of the first version, on a sample six times larger. On two files that also have a published edition the estimate can be checked against ground truth: it is unbiased to within a percentage point and ranks pages as the truth does (rank correlation 0.92 where the edition is a faithful witness). This bounds use: two variants of one work here share 19% of words, below the rate at which two readings of a single page agree, so the redactions cannot be collated word by word at this quality. That negative result is reported as such, and the constraint is built into the tool.

cs.CL↗

How Far Do On-Prem Open LLMs Get on Text-to-SQL? A Cross-Family Size x Technique Frontier on BIRD

Organizations that cannot send data to a cloud API increasingly ask: how good is Text-to-SQL if the model must run on-premises on open weights, and which popular accuracy "recipes" are worth their compute? We answer with an honest, fully reproducible benchmark on the BIRD development split (n=1534, Execution Accuracy), evaluating three open model families across two generations -- Qwen2.5-Coder (7B/14B/32B), CodeLlama-Instruct (7B/13B/34B), and Llama-3.x (8B, 70B) -- under one matched protocol, ablating a model-agnostic recipe (schema linking, self-correction, self-consistency) component by component, with every difference tested by the paired McNemar test. Four findings stand out. (i) Generation matters more than raw size, and the recipe is family-robust: Qwen2.5-Coder dominates the older CodeLlama at matched size (39.1 vs 20.9 at 7B), but a modern non-Qwen model (Llama-3.3-70B, 49.2 on a matched serving) is competitive, so CodeLlama's weakness reflects its 2023 generation, not "non-Qwen = weak". (ii) Self-correction is a robust, near-free win, significant on all three families where there is room to improve. (iii) Schema linking does not help, and a stronger linker does not rescue it: a retrieval/embedding linker with 96.5% gold-table recall is statistically indistinguishable from no linking, ruling out the "weak lexical strawman" objection across three families. (iv) Self-consistency is poor value (+0.13 pp for ~5x tokens, not significant). We report real per-stage cost ($/1k queries) and release all code, predictions, and summaries; archived code and data: https://doi.org/10.5281/zenodo.20952794

cs.CL↗

AI Appeals Processor: A Deep Learning Approach to Automated Classification of Citizen Appeals in Government Services

Government agencies worldwide face growing volumes of citizen appeals, with electronic submissions increasing significantly over recent years. Traditional manual processing averages 20 minutes per appeal with only 67% classification accuracy, creating significant bottlenecks in public service delivery. This paper presents AI Appeals Processor, a microservice-based system that integrates natural language processing and deep learning techniques for automated classification and routing of citizen appeals. We evaluate multiple approaches -- including Bag-of-Words with SVM, TF-IDF with SVM, fastText, Word2Vec with LSTM, and BERT -- on a representative dataset of 10,000 real citizen appeals across three primary categories (complaints, applications, and proposals) and seven thematic domains. Our experiments demonstrate that a Word2Vec+LSTM architecture achieves 78% classification accuracy while reducing processing time by 54%, offering an optimal balance between accuracy and computational efficiency compared to transformer-based models.

cs.CL↗