arXiv Science⌕ Search

arXiv · 0705.4442

World-set Decompositions: Expressiveness and Efficient Algorithms

Abstract

Uncertain information is commonplace in real-world data management scenarios. The ability to represent large sets of possible instances (worlds) while supporting efficient storage and processing is an important challenge in this context. The recent formalism of world-set decompositions (WSDs) provides a space-efficient representation for uncertain data that also supports scalable processing. WSDs are complete for finite world-sets in that they can represent any finite set of possible worlds. For possibly infinite world-sets, we show that a natural generalization of WSDs precisely captures the expressive power of c-tables. We then show that several important decision problems are efficiently solvable on WSDs while they are NP-hard on c-tables. Finally, we give a polynomial-time algorithm for factorizing WSDs, i.e. an efficient algorithm for minimizing such representations.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Dan Olteanu, Christoph Koch, Lyublena Antova. 2008-01-09. World-set Decompositions: Expressiveness and Efficient Algorithms. https://arxiv.org/abs/0705.4442

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

IBLTs Measure Before They Decode: Self-Sizing Set Reconciliation for Database Consistency Verification

Cross-system data replication pipelines cannot confirm end-to-end consistency from the local guarantees of each hop, so the two endpoints must be compared directly on a periodic basis. Once the rows of a fixed snapshot are normalized into fingerprints, the task reduces to finding the symmetric difference of the two sets. An Invertible Bloom Lookup Table (IBLT) reconciles the sets with communication that grows only with the difference cardinality $d$, independent of table size, but its capacity must be fixed while $d$ is still unknown. Across 41,603 production reconciliations over 90 days, nonzero $d$ spans about seven orders of magnitude, and no reliable empirical constant exists. We show that the count array of an IBLT has already measured $d$ before decoding. The measurement is in-band: it is carried by the recovery sketch itself and adds no bytes dedicated to estimation. A mapping-aware theorem extends the construction to Irregular, Rateless, and MET IBLTs. The protocol reads the estimate only after a decoding failure; we prove that the failure-conditioned lower quantile bounds the risk of underestimation, which gives the second-round capacity a configurable success-probability guarantee. The resulting self-sizing protocol attempts recovery with a small first-round sketch and stops on success; on failure it reads $d$, sizes the second round, and completes reconciliation in at most two rounds. Against a controlled oracle, communication is 1.29-1.47 times that of a scheme given $d$ in advance. Production workload characterization, relational-database replay, and a cross-city KV deployment confirm the end-to-end mechanism. In production on an Oracle-MySQL link, all completed runs succeeded within two rounds, over 90% on the 1-RTT fast path with a single 16 KB sketch.

cs.DB↗

A 3GPP-Compliant Benchmark Dataset for RIS-Aided Beyond 5G Networks

Reconfigurable Intelligent Surfaces (RIS) are emerging as a key technology for programmable wireless environments in the beyond the fifth generation (B5G) networks. However, data-driven RIS research remains bottleneck by the lack of standardized, high-fidelity and open-source datasets. In this paper, we introduce a large-scale 3GPP TR 38.901-compliant dataset for RIS-aided millimeter wave (mmWave) networks, that considers severe path loss, blockage sensitivity, and spatial channel sparsity make the RIS assistance more impactful. The dataset spans various canonical 3GPP deployment scenarios across 20 controlled variants, capturing diverse user densities, fading conditions, and blockage regimes. Uniquely, every sample includes oracle RIS phase configurations obtained via a globally optimal brute-force codebook search, providing gold-standard supervision labels that are absent from any existing public dataset. Rich multi-task annotations comprising full channel state information (CSI), per-link channel decomposition, optimal phase matrices, and channel quality index (CQI) labels support a broad range of machine learning paradigms and downstream tasks, including phase optimization, channel estimation, and interference management. As the primary benchmark task, we introduce a novel CSI-to-CQI mapping that frames RIS-aided link-quality prediction as a scalable scalar classification problem, thereby avoiding the exponential output complexity of the direct phase vector prediction. We have evaluated this mapping against state-of-the-art architectures under in-distribution, out-of-distribution, and real-world hardware measurement conditions. Our dataset provides a reproducible, extensible, and community-ready foundation to accelerate data-driven research in RIS-aided B5G networks.

cs.DB↗

X-DigCheck: Co-Evolving Application Profiles and Knowledge Graphs, Demonstrated on the RTI Documentation of Rupe Magna

We demonstrate X-DigCheck, a domain-independent environment for building and maintaining application profiles as they co-evolve with the data they describe. Profiles developed against a fixed ontology quickly drift from the schema they were meant to capture. X-DigCheck treats profile construction as a continuous ontology-data co-evolution loop: data are lifted into RDF against the profile, checked through competency questions and SHACL, and the resulting reports jointly drive revisions of the ontology, mappings, constraints, and graph. The loop is agnostic to the domain and to the pipeline that produces the graph. We validate and demonstrate the tool in the cultural heritage domain, on the construction of RupeMagna-RTI, the first Reflectance Transformation Imaging (RTI) specialisation of the Cultural Heritage Survey ODP (CHS-ODP), aligned with CIDOC-CRM/CRMdig, ArCo, CHAD-KG, and Getty AAT, with semRTI as the lifting pipeline of this use case. The demonstration lets visitors run one full turn of the loop -on the shipped Rupe Magna (Grosio, Italy) RTI survey, or on a profile and graph of their own -executing the competency-question and SHACL checks live and reading the bidirectional coverage report that flags modelling gaps and stale assumptions. The result is a portable co-evolution environment for profile engineering, together with a reusable RTI application profile produced through it. A screencast of the demonstration is available at https://zenodo.org/records/22210609.

cs.DB↗