arXiv ScienceSearch

arXiv subjects

Alex Smolin

Publications and source records attributed to Alex Smolin.

9 recordsLinked to original sources

Beliefs and Behavior in Language Models

There is significant uncertainty about whether abstractions like beliefs or desires usefully describe the behavior of large language models (LLMs). In addition to the inherent scientific interest of this question, these latent quantities are often invoked to explain the behavior of LLMs to users or to define and evaluate harmful behaviors which are relative to intent. Nevertheless, we currently lack a means to systematically test whether concepts like "belief" are well-applied to LLMs, and hence whether they are likely to be fruitful ingredients of attempts to align models with human interests. We propose an approach for empirically studying such questions, asking whether a single latent variable inferred from the LLMs' outputs -- interpreted as a degree of belief -- allows an observer to make interpretable predictions of how the LLMs' will respond to new prompts. We find that highly capable models are usefully described as holding beliefs and that, generally, the predictability of model outputs based on an inferred latent belief tracks overall trends in model capability. Building on these findings, we provide empirical strategies to study how beliefs in LLMs can be measured, the extent to which LLMs comply with instructed decision rules or payoffs, and how beliefs evolve within individual instances of an LLM over the course of reasoning.

cs.AI

Robust Trust

An agent chooses an action based on her private information and a recommendation from an informed but potentially misaligned adviser. With a known probability, the adviser truthfully reports his signal; with the remaining probability, he can send any message. We characterize optimal robust decision rules that maximize the agent's worst-case expected payoff. Every optimal rule is equivalent to a trust-region policy in belief space: the adviser's reported beliefs are taken at face value if they fall within the trust region but are otherwise clipped to the trust region's boundary. We derive alignment thresholds above which advice is strictly valuable and fully characterize the solution in both binary-state and binary-action environments.

econ.TH

Calibrated Mechanism Design

We study mechanism design when a designer repeatedly uses a fixed mechanism to interact with strategic agents who learn from observing their allocations. We introduce a static framework, calibrated mechanism design, requiring mechanisms to remain incentive compatible given the information they reveal about an underlying state through repeated use. In single-agent settings, we prove implementable outcomes correspond to two-stage mechanisms: the designer discloses information about the state, then commits to a state-independent allocation rule. This yields a tractable procedure to characterize calibrated mechanisms, combining information design and mechanism design. In private values environments, full transparency is optimal and correlation-based surplus extraction fails. We provide a microfoundation by showing calibrated mechanisms characterize exactly what is implementable when an infinitely patient agent repeatedly interacts with the same mechanism. Dynamic mechanisms that condition on histories expand implementable outcomes only by weakening incentive constraints, but not by enriching the designer's ability to obfuscate learning.

econ.TH

Menu Pricing of Large Language Models

We develop a framework for the optimal pricing and product design of LLMs in which a provider sells menus of token budgets to users who differ in their valuations across a continuum of tasks. Under a homogeneous production technology, we show that users' high-dimensional type profiles are summarized by a scalar index, reducing the seller's problem to one-dimensional screening. The optimal mechanism takes the form of committed-spend contracts: buyers pay for a budget that they allocate across token classes priced at marginal cost. We extend the analysis to environments with multiple differentiated models and to competition between a proprietary leader and an open-source fringe, showing that competitive pressure reshapes both the intensive and extensive margins of compute provision. Each element of our theory (token-budget menus, maximum- and minimum-spend plans, multi-model versioning, and linear API pricing) has a direct counterpart in the observed pricing practices of providers such as Anthropic, OpenAI, and GitHub.

econ.TH

Algorithmic Recommendations and Strategic Pricing

We study algorithmic recommendations in a market with a buyer, privately informed sellers, and an algorithm that can recommend a product based on product values and posted prices but cannot use transfers or force purchases. For objectives ranging from total surplus to buyer surplus, algorithmic recommendations implement the same welfare outcomes as direct mechanisms. We study the effects of algorithmic recommendations on pricing, competition, the composition of trade, and market segmentation. Our results inform the operation of AI assistants and other technologies that mediate product discovery, attention, and pricing.

econ.TH

Data Provision to an Informed Seller

A monopoly seller is privately and imperfectly informed about the buyer's value of the product. The seller uses information to price discriminate the buyer. A designer offers a mechanism that provides the seller with additional information based on the seller's report about her type. We establish the impossibility of screening for welfare purposes, i.e., the designer can attain any implementable combination of buyer surplus and seller profit by providing the same signal to all seller types. We use this result to characterize the set of implementable welfare outcomes, study the seller's incentive to acquire third-party data, and demonstrate the trade-off between buyer surplus and efficiency.

econ.TH

Information Design in Smooth Games

We study information design in games where players choose from a continuum of actions and have continuously differentiable payoffs. We show that an information structure is optimal when the equilibrium it induces can also be implemented in a principal-agent contracting problem. Building on this result, we characterize optimal information structures in symmetric linear-quadratic games. With common values, targeted disclosure is robustly optimal across all priors. With interdependent and normally distributed values, linear disclosure is uniquely optimal. We illustrate our findings with applications in venture capital, Bayesian polarization, and price competition.

econ.TH

Persuasion and Welfare

Information policies such as scores, ratings, and recommendations are increasingly shaping society's choices in high-stakes domains. We provide a framework to study the welfare implications of information policies on a population of heterogeneous individuals. We define and characterize the Bayes welfare set, consisting of the population's utility profiles that are feasible under some information policy. The Pareto frontier of this set can be recovered by a series of standard Bayesian persuasion problems, in which a utilitarian planner takes the role of the information designer. We provide necessary and sufficient conditions under which an information policy exists that Pareto dominates the no-information policy. We illustrate our results with applications to data leakage, price discrimination, and credit ratings.

econ.TH

The Optimality of Upgrade Pricing

We consider a multiproduct monopoly pricing model. We provide sufficient conditions under which the optimal mechanism can be implemented via upgrade pricing -- a menu of product bundles that are nested in the strong set order. Our approach exploits duality methods to identify conditions on the distribution of consumer types under which (a) each product is purchased by the same set of buyers as under separate monopoly pricing (though the transfers can be different), and (b) these sets are nested. We exhibit two distinct sets of sufficient conditions. The first set of conditions is given by a weak version of monotonicity of types and virtual values, while maintaining a regularity assumption, i.e., that the product-by-product revenue curves are single-peaked. The second set of conditions establishes the optimality of upgrade pricing for type spaces with monotone marginal rates of substitution (MRS) -- the relative preference ratios for any two products are monotone across types. The monotone MRS condition allows us to relax the earlier regularity assumption. Under both sets of conditions, we fully characterize the product bundles and prices that form the optimal upgrade pricing menu. Finally, we show that, if the consumer's types are monotone, the seller can equivalently post a vector of single-item prices: upgrade pricing and separate pricing are equivalent.

cs.GT