arXiv ScienceSearch

arXiv subjects

Ohad Perry

Publications and source records attributed to Ohad Perry.

12 recordsLinked to original sources

Stability of Fork-Join Systems with Redundancy and Heterogeneous Servers

We consider the stability problem of fork-join systems with redundancy (FJR) and heterogeneous servers under both static and dynamic capacity-allocation policies. In an $(n,k)$ FJR system, each arriving job is split into $n$ independent tasks, with one task assigned to each of $n$ parallel servers. Once $k \le n$ tasks have been processed, they are joined and the corresponding job departs the system; the remaining $n-k$ unprocessed tasks are then removed and are therefore termed redundant. We first identify the nominal traffic intensity and characterize the maximal stability region, defined as the set of traffic intensities for which there exists an admissible policy that stabilizes the system. We then establish conditions under which this maximal stability region is attained for two classes of policies: static and dynamic. Specifically, we show that for static allocation policies, in which service capacities remain fixed over time, maximality is achieved whenever the fastest server is allocated no more than $1/k$ of the total service capacity. For dynamic allocation policies, in which a fixed total service capacity may be repeatedly reallocated among the servers, we show that maximality is achieved whenever the cumulative capacity allocated to the $j$ shortest queues does not exceed $j/k$ of the total capacity for every $j=1,\ldots,k-1$. Our analysis is based on a projection of the $(n,k)$ FJR system onto a simpler $(k,k)$ system that has no redundancy, together with a novel sample-path comparison argument for multidimensional processes based on the generalized Schur-convex order.

cs.IT

A Queueing Model of Patient Flow for Stroke Networks to Estimate Acute Stroke Transfer Capacity

Background: Most acute stroke (AS) patients in the United States are initially evaluated at a primary stroke center (PSC) and a significant proportion requires transfer to a comprehensive stroke center (CSC) for advanced treatment. A CSC typically accepts patients from multiple PSCs in its network, leading to capacity limits. This study uses a queueing model to estimate impacts on CSC capacity due to transfers from PSCs. Methods: The model assumes that the number of AS patients arriving at each PSC, proportion of AS patients transferred, and length of stay in the CSC Neurologic Intensive Care Unit (Neuro-ICU) by type of AS are random, while the transfer rates of ischemic and hemorrhagic AS patients are control variables. The main outcome measure is the "overflow" probability, namely, the probability of a CSC not having capacity (unavailability of a Neuro-ICU bed) to accept a transfer. Data simulations of the model, using a base case and an expanded case, were performed to illustrate the effects of changing key parameters, such as transfer rates from PSCs and CSC Neuro-ICU capacity on overflow capacity. Results: Data simulations of the model using a base case show that an increase of a PSC's ischemic stroke transfer rate from 15% to 55% raises the overflow probability from 30.62% to 36.13%. Further simulations of the expanded case show that to maintain an a priori CSC overflow probability of 30.62% when adding a PSC with a AS transfer rate of 15% to the network, other PSCs would need to decrease their transfer rate by 12.5% or the CSC Neuro-ICU would need to add 2 beds. Discussion: A queuing model can be used to estimate the effects of change in the size of a PSC-CSC network, change in AS transfer rates, or change in number of CSC Neuro-ICU beds of a CSC on its capacity on the overflow probability in the CSC.

q-bio.QM

Many-Server Heavy-Traffic Limits for Queueing Systems with Perfectly Correlated Service and Patience Times

We characterize heavy-traffic process and steady-state limits for systems staffed according to the square-root safety rule, when the service requirements of the customers are perfectly correlated with their individual patience for waiting in queue. Under the usual many-server diffusion scaling, we show that the system is asymptotically equivalent to a system with no abandonment. In particular, the limit is the Halfin-Whitt diffusion for the $M/M/n$ queue when the traffic intensity approaches its critical value $1$ from below, and is otherwise a transient diffusion, despite the fact that the prelimit is positive recurrent. To obtain a refined measure of the congestion due to the correlation, we characterize a lower-order fluid (LOF) limit for the case in which the diffusion limit is transient, demonstrating that the queue in this case scales like $n^{3/4}$. Under both the diffusion and LOF scalings, we show that the stationary distributions converge weakly to the time-limiting behavior of the corresponding process limit.

math.PR

Existence and Approximations of Moments for Polling Systems under the Binomial-Exhaustive Policy

We establish sufficient conditions for the existence of moments of the steady-state queue in polling systems operating under the binomial-exhaustive policy (BEP). We assume that the server switches between the different buffers according to a pre-specified table, and that switchover times are incurred whenever the server moves from one buffer to the next. We further assume that customers arrive according to independent Poisson processes, and that the service and switchover times are independent random variables with general distributions. We then propose a simple scheme to approximate the moments, which is shown to be asymptotically exact as the switchover times grow without bound, and whose computation complexity does not grow with the order of the moment. Finally, we demonstrate that the proposed asymptotic approximation for the moments is related to the fluid limit under a large-switchover-time scaling; thus, similar approximations can be easily derived for other server-switching policies, by simply identifying the fluid limits under those controls. Numerical examples demonstrate the effectiveness of our approximations for the moments under BEP and under other policies, and their increased accuracy as the switchover times increase.

math.PR

Asymptotic Optimality of the Binomial-Exhaustive Policy for Polling Systems with Large Switchover Times

We study an optimal-control problem of polling systems with large switchover times, when a holding cost is incurred on the queues. In particular, we consider a stochastic network with a single server that switches between several buffers (queues) according to a pre-specified order, assuming that the switchover times between the queues are large relative to the processing times of individual jobs. Due to its complexity, computing an optimal control for such a system is prohibitive, and so we instead search for an asymptotically optimal control. To this end, we first solve an optimal control problem for a deterministic relaxation (namely, for a fluid model), that is represented as a hybrid dynamical system. We then "translate" the solution to that fluid problem to a binomial-exhaustive policy for the underlying stochastic system, and prove that this policy is asymptotically optimal in a large-switchover-time scaling regime, provided a certain uniform integrability (UI) condition holds. Finally, we demonstrate that the aforementioned UI condition holds in the following cases: (i) the holding cost has (at most) linear growth, and all service times have finite second moments; (ii) the holding cost grows at most at a polynomial rate (of any degree), and the service-time distributions possess finite moment generating functions.

math.PR

Stability of Parallel Server Systems

The fundamental problem in the study of parallel-server systems is that of finding and analyzing `good' routing policies of arriving jobs to the servers. It is well known that, if full information regarding the workload process is available to a central dispatcher, then the {\em join the shortest workload} (JSW) policy, which assigns jobs to the server with the least workload, is the optimal assignment policy, in that it maximizes server utilization, and thus minimizes sojourn times. The {\em join the shortest queue} (JSQ) policy is an efficient dispatching policy when information is available only on the number of jobs with each of the servers, but not on their service requirements. If information on the state of the system is not available, other dispatching policies need to be employed, such as the power-of-$d$ routing policy, in which each arriving job joins the shortest among $d \ge 1$ queues sampled uniformly at random. (Under this latter policy, the system is known as {\em the supermarket model}.) In this paper we study the stability question of parallel server systems assuming that routing errors occur, so that arrivals may be routed to the `wrong' (not to the smallest) queue with a positive probability. We show that, even if a `non-idling' dispatching policy is employed, under which new arrivals are always routed to an idle server, if any is available, the performance of the system can be much worse than under the policy that chooses one of the servers uniformly at random. More specifically, we prove that the usual traffic intensity $\rho < 1$ does not guarantee that the system is stable.

math.PR

On the Instability of Matching Queues

A matching queue is described via a graph $G$ together with a matching policy. Specifically, to each node in the graph there is a corresponding arrival process of items which can either be queued, or matched with queued items in neighboring nodes. The matching policy specifies how items are matched whenever more than one matching is possible. Motivated by the increasing theoretical interest in such matching models, we investigate the question of (in)stability of matching queues which satisfy a natural necessary condition for stability, which can be thought of as an analogue of the usual traffic condition for traditional queueing networks (namely, $\rho_i < 1$ in each service station $i$). We employ fluid-stability arguments to show that matching queues can in general be unstable, even though the necessary stability condition is satisfied.

math.PR

A Switching Fluid Limit of a Stochastic Network Under a State-Space-Collapse Inducing Control with Chattering

Routing mechanisms for stochastic networks are often designed to produce state space collapse (SSC) in a heavy-traffic limit, i.e., to confine the limiting process to a lower-dimensional subset of its full state space. In a fluid limit, a control producing asymptotic SSC corresponds to an ideal sliding mode control that forces the fluid trajectories to a lower-dimensional sliding manifold. Within deterministic dynamical systems theory, it is well known that sliding-mode controls can cause the system to chatter back and forth along the sliding manifold due to delays in activation of the control. For the prelimit stochastic system, chattering implies fluid-scaled fluctuations that are larger than typical stochastic fluctuations. In this paper we show that chattering can occur in the fluid limit of a controlled stochastic network when inappropriate control parameters are used. The model has two large service pools operating under the fixed-queue-ratio with activation and release thresholds (FQR-ART) overload control which we proposed in a recent paper. We now show that, if the control parameters are not chosen properly, then delays in activating and releasing the control can cause chattering with large oscillations in the fluid limit. In turn, these fluid-scaled fluctuations lead to severe congestion, even when the arrival rates are smaller than the potential total service rate in the system, a phenomenon referred to as congestion collapse. We show that the fluid limit can be a bi-stable switching system possessing a unique nontrivial periodic equilibrium, in addition to a unique stationary point.

math.PR

Achieving Rapid Recovery in an Overload Control for Large-Scale Service Systems

We consider an automatic overload control for two large service systems modeled as multi-server queues, such as call centers. We assume that the two systems are designed to operate independently, but want to help each other respond to unexpected overloads. The proposed overload control automatically activates sharing (sending some customers from one system to the other) once a ratio of the queue lengths in the two systems crosses an activation threshold (with ratio and activation threshold parameters for each direction). To prevent harmful sharing, sharing is allowed in only one direction at any time. In this paper, we are primarily concerned with ensuring that the system recovers rapidly after the overload is over, either (i) because the two systems return to normal loading or (ii) because the direction of the overload suddenly shifts in the opposite direction. To achieve rapid recovery, we introduce lower thresholds for the queue ratios, below which one-way sharing is released. As a basis for studying the complex dynamics, we develop a new six-dimensional fluid approximation for a system with time-varying arrival rates, extending a previous fluid approximation involving a stochastic averaging principle. We conduct simulations to confirm that the new algorithm is effective for predicting the system performance and choosing effective control parameters. The simulation and the algorithm both show that the system can experience an inefficient nearly-periodic behavior, corresponding to an oscillating equilibrium (congestion collapse), if the sharing is strongly inefficient and the control parameters are set inappropriately.

math.PR

Diffusion Approximation for an Overloaded X Model Via a Stochastic Averaging Principle

In previous papers we developed a deterministic fluid approximation for an overloaded Markovian queueing system having two customer classes and two service pools, known in the call-center literature as the X model. The system uses the fixed-queue-ratio-with-thresholds (FQR-T) control, which we proposed in a recent paper as a way for one service system to help another in face of an unexpected overload. Under FQR-T, customers are served by their own service pool until a threshold is exceeded. Then, one-way sharing is activated with customers from one class allowed to be served in both pools. The control aims to keep the two queues at a pre-specified fixed ratio. We supported the fluid approximation by establishing a many-server heavy-traffic functional weak law of large numbers (FWLLN) involving an averaging principle. In this paper we develop a refined diffusion approximation for the same model based on a many-server heavy-traffic functional central limit theorem (FCLT).

math.PR

An ODE for an Overloaded X Model Involving a Stochastic Averaging Principle

We study an ordinary differential equation (ODE) arising as the many-server heavy-traffic fluid limit of a sequence of overloaded Markovian queueing models with two customer classes and two service pools. The system, known as the X model in the call-center literature, operates under the fixed-queue-ratio-with-thresholds (FQR-T) control, which we proposed in a recent paper as a way for one service system to help another in face of an unanticipated overload. Each pool serves only its own class until a threshold is exceeded; then one-way sharing is activated with all customer-server assignments then driving the two queues toward a fixed ratio. For large systems, that fixed ratio is achieved approximately. The ODE describes system performance during an overload. The control is driven by a queue-difference stochastic process, which operates in a faster time scale than the queueing processes themselves, thus achieving a time-dependent steady state instantaneously in the limit. As a result, for the ODE, the driving process is replaced by its long-run average behavior at each instant of time; i.e., the ODE involves a heavy-traffic averaging principle (AP).

math.PR

A Fluid Limit for an Overloaded X Model Via a Stochastic Averaging Principle

We prove a many-server heavy-traffic fluid limit for an overloaded Markovian queueing system having two customer classes and two service pools, known in the call-center literature as the X model. The system uses the fixed-queue-ratio-with-thresholds (FQR-T) control, which we proposed in a recent paper as a way for one service system to help another in face of an unexpected overload. Under FQR-T, customers are served by their own service pool until a threshold is exceeded. Then, one-way sharing is activated with customers from one class allowed to be served in both pools. After the control is activated, it aims to keep the two queues at a pre-specified fixed ratio. For large systems that fixed ratio is achieved approximately. For the fluid limit, or FWLLN, we consider a sequence of properly scaled X models in overload operating under FQR-T. Our proof of the FWLLN follows the compactness approach, i.e., we show that the sequence of scaled processes is tight, and then show that all converging subsequences have the specified limit. The characterization step is complicated because the queue-difference processes, which determine the customer-server assignments, remain stochastically bounded, and need to be considered without spatial scaling. Asymptotically, these queue-difference processes operate in a faster time scale than the fluid-scaled processes. In the limit, due to a separation of time scales, the driving processes converge to a time-dependent steady state (or local average) of a time-varying fast-time-scale process (FTSP). This averaging principle (AP) allows us to replace the driving processes with the long-run average behavior of the FTSP.

math.PR