with-limits keeps a development machine responsive while background commands
run. It gives foreground applications scheduling priority, limits the child
process tree's memory, sustained CPU use, and wall-clock runtime, and works on
macOS, Windows, and Linux.
The motivating case is concurrent agents doing machine-learning or research work. Several agents can each start a RAM-intensive job that looks reasonable in isolation, while their combined memory use exhausts and crashes the host. Putting a process-tree limit around each job contains that failure.
The entire child tree also runs at a fixed lower scheduling priority. On Unix, its inherited niceness is increased by 10, capped at the maximum of 19. Windows uses Below Normal priority. This still lets a job use otherwise-idle CPU capacity while allowing interactive applications to win when they need it.
For a time limit alone, GNU timeout (installed as gtimeout by Homebrew) is
the established choice. with-limits is useful when one portable command
should also constrain memory or sustained CPU consumption.
with-limits is an early-stage, actively maintained tool. Releases before 1.0
may change command-line options, environment variables, and exit-status
behavior.
Installation requires Rust 1.85 or newer and Cargo. The Rust toolchain installer provides both.
Install from crates.io:
cargo install with-limits --lockedThis installs with-limits in Cargo's binary directory, normally
~/.cargo/bin. Ensure that directory is on your PATH.
git clone https://github.com/osteele/with-limits.git
cd with-limits
cargo install --path .with-limits is intended for trusted local and CI workloads. It is a resource
guard, not a security sandbox for hostile code.
Verify the installation with a command available alongside Cargo:
with-limits --time 5s -- cargo --versionA successful run prints the Cargo version and exits with status zero. In an interactive terminal, the startup summary also reports the time limit and lower scheduling priority.
Place the command after --:
with-limits --memory 8GiB --cpu 2 --time 30m -- python train.pyUse -c to run a shell command. On macOS this uses /bin/zsh; on other Unix
systems it uses /bin/sh; and on Windows it uses %COMSPEC%, falling back to
cmd.exe when that variable is unset.
with-limits --memory 4GiB -c 'just format && just check'When no limit option is supplied, with-limits applies --memory auto, which
is 70% of the memory available when the command starts. It also preserves the
remaining 30% as a continuously checked host reserve. This lets several
concurrent guarded agents react to their combined memory use:
with-limits -c 'just format && just check'Lower scheduling priority is independent of the resource-limit options and is enabled for every command by default.
Options:
--memory SIZE,-m SIZElimits aggregate resident memory. Sizes accept SI suffixes such asGB, IEC suffixes such asGiB,auto, or a percentage of initially available memory such as60%. On macOS, available memory is the kernel's free percentage (kern.memorystatus_level) of physical memory. A percentage orautolimit that resolves below 256 MiB exits75without running the command.--reserve SIZEpublishes an absolute memory estimate for admission, independently of the containment cap. An explicit absolute--memorysize is also the reservation when this option is absent.--cpu CORESlimits sustained CPU use, where1is the capacity of one logical core. Fractional values such as0.5are accepted; the minimum is0.01.--time DURATION,-t DURATIONlimits wall-clock runtime. Durations use the GNUtimeoutsuffixess,m,h, andd; milliseconds usems.--kill-after DURATION(also--grace) controls how long graceful termination may take before the process tree is forcibly terminated. The default is two seconds.--shell PATHselects the shell used by-c.--require-nativerejects a requested memory or CPU limit if the platform would enforce it by sampling. Memory limits are currently sampled on every platform; Windows CPU limits are native.--foregroundkeeps the command in the supervisor's process group so it can read from the terminal. See Enforcement for the tradeoff. Unix only; accepted and ignored on Windows.--quiet,-qsuppresses the interactive startup summary. Limit violations are still reported.--no-reservationignores outstanding admission reservations and does not publish one for the command.
Containment limits a command after it starts. The admission gate answers the earlier question of whether to start one at all, so several independent agents can consult the same host policy instead of each keeping its own copy. The resource that exhausts first on a machine running many agents is RAM, so the gate is built on memory signals: CPU load is the symptom of paging, and a load threshold fires late and on the wrong quantity.
Four signals feed the decision, and two of them are enforced by default.
The kernel memory pressure level (1 normal, 2 warn, 4 critical) is read on
macOS from kern.memorystatus_vm_pressure_level and on Linux derived from the
PSI /proc/pressure/memory stall time — a mapping onto the macOS scale chosen
for this tool, not a kernel verdict; it is unknown on Windows. The default
refuses only critical. Warn is the ordinary operating state of a machine
running many concurrent tasks: measured on one such workstation it persisted
for hours at a stretch while a third of memory stayed free, so refusing warn
parks work indefinitely without reducing risk.
Free memory as a percentage of total is the signal that separates ordinary
memory management from exhaustion, and the default floor is 10%. On macOS it
is the kernel's own kern.memorystatus_level, the figure Apple's
memory_pressure tool prints as the free percentage. It counts every page
that is neither wired nor held by the compressor, active pages included, so
it is well above the vm_stat free count: a host with 0.4% of pages free and
40% in the compressor reported 31%. Elsewhere it is derived from available and
total memory. On the workstation above, every state the gate
refused sat between 27% and 53% free, while the state it exists to refuse had
32 MB free, effectively zero.
Swap free and swap total are read on every platform and are reported, not
enforced, by default. Free space inside the current swap files does not
measure exhaustion on macOS: swap is sized on demand and grows while free disk
allows, so the fraction in use rises simply because the kernel sized swap to
the workload. Setting --min-swap-free enforces a floor for a platform where
swap is a fixed partition; a swap total at or below that floor is then exempt,
since swap that small has barely been touched.
The one-minute load average per logical CPU is read on macOS and Linux and is unknown on Windows; it is reported but not enforced unless a ceiling is set. CPU load is the late symptom of paging, not the resource that runs out.
Every signal is optional. A platform that cannot answer, a failed sysctl, or
an unparsable file yields an unknown for that signal only, and unknown is
never replaced by a fabricated value. An unknown signal is reported by name
but does not refuse, because whether "cannot tell" is a refusal is the
caller's policy; --refuse-unknown treats an unknown enforced signal as a
refusal for the caller that wants that.
An explicit absolute reservation publishes an admission promise. For example,
--memory auto --reserve 3GiB allows the tree to use the automatic containment
cap while asking admission checks to assume about 3 GiB of growth. Other
admission checks subtract the unrealized part of that promise from free memory.
The unrealized amount is the reservation minus the tree's most recently
sampled RSS, with a floor of zero. A tree that grows past its reservation
contributes no unrealized bytes but remains governed by its independent memory
cap. An observation older than 30 seconds contributes the full reservation.
The supervisor refreshes and retains its reservation while the store remains available. A refresh never waits for the store lock: when another process holds it, that tick's refresh is skipped. A refresh failure prints one warning, removes the reservation, and continues supervising and terminating the command. Each record includes the supervisor's process start time so a recycled process id does not keep an orphaned reservation alive. Platforms that cannot read a start time fall back to process-id liveness.
The admission decision and reservation write share one exclusive file lock. Two callers that reach the gate together therefore make their decisions in sequence, and the second sees the first caller's promise. The lock is held only while reading, deciding, and writing. It is released before the command or a headroom wait begins.
auto and percentage memory limits do not implicitly publish reservations.
They cap a tree relative to whatever memory is free, so using the cap as the
promise would let the first caller claim most of the host and serialize later
work. Pair either form with an absolute estimate when admission coordination is
required:
with-limits --wait-for-headroom=15m --memory auto --reserve 3GiB -- commandAn absolute --memory size remains an implicit reservation when --reserve is
absent. --reserve is valid without --memory. When it is the only resource
option, the usual default --memory auto containment applies. If --cpu or
--time is also present, omitting --memory leaves memory uncapped.
--check-headroom reads reservations but never creates one.
with-limits --check-headroom reads the signals and outstanding reservations,
applies the policy, prints one line per signal plus the reservation count and
unrealized byte total, and exits 0 when admitted or 75 (EX_TEMPFAIL) when
refused. --json prints the readings, the effective policy, the reservation
summary, the refusing reasons, and the unknown signals as one JSON object
instead:
with-limits --check-headroom
with-limits --check-headroom --jsonwith-limits --wait-for-headroom[=DURATION] [limits] -- command polls the
same decision until admitted, then runs the command under the given limits.
The wait is unbounded without a DURATION; with one, expiry exits 75 without
running the command. Polls back off from 15 seconds by a factor of 1.5 with a
random multiplier in [0.5, 1.5), capped at 90 seconds per sleep. The jitter
matters because dozens of agent sessions poll independently: a fixed backoff
makes them retry in lockstep and start together the moment pressure clears.
While waiting, a line naming the refusing signals goes to stderr on the first
refusal and at most once a minute, unless --quiet. A terminating signal
during the wait ends it, since no command is running yet.
The wait never implies a limit and a limit never implies a wait; with no
limit options the default --memory auto applies as usual.
--max-pressure Nsets the highest admitted pressure level. The default is2, so only critical pressure refuses.--min-memory-free-percent PERCENTsets the free-memory floor. The default is10;0reports free memory without enforcing it.--min-swap-free SIZEsets the swap floor as an absolute size (percentages are not accepted). The default is0, which reports swap without enforcing it.--max-load-per-cpu CORESenforces a load ceiling. Without it, load is reported but never refuses.--refuse-unknowntreats an unknown enforced signal as a refusal.--reserve SIZEsets the absolute estimate published when the command is admitted. It overrides the implicit reservation from an absolute--memorysize and is valid without an explicit memory cap.--no-reservationbypasses the reservation store for this invocation. It neither subtracts other reservations nor publishes its own.
Agent lifecycle hooks can place selected shell workloads under with-limits
without changing the command the model generates. The included example wraps
POSIX uv run requests, including the rest of a pipeline or compound command,
under one default memory budget:
uv run python experiment.py && just process-results
→ with-limits -c 'uv run python experiment.py && just process-results'
| Agent | Automatic wrapping path |
|---|---|
| Claude Code | Native PreToolUse command hook |
| Codex | Native PreToolUse command hook |
| OpenCode | tool.execute.before plugin |
| Kimi Code CLI | Session-wide launcher or command-specific PATH wrapper |
Claude Code and Codex can use the same hook script. OpenCode reaches it through a small plugin. Kimi hooks cannot currently replace a tool input, so they cannot apply this command-by-command rewrite. See Agent shell hooks for the example, configuration, security tradeoff, and platform scope.
WITH_LIMITS_NICE controls fixed priority lowering. It defaults to 10. On
Unix, values from 1 through 19 are added to the child process's inherited
niceness, capped at 19; descendants inherit the result. On Windows, any enabled
value selects the Below Normal priority class for the Job Object and therefore
its entire process tree. Set it to 0, off, false, or no to disable
priority lowering.
The priority does not change in response to load or the number of running agents. Fixed lower priority lets foreground work preempt background jobs while leaving idle CPU capacity available to them.
On macOS, guarded commands receive PYTORCH_MPS_HIGH_WATERMARK_RATIO (0.7)
and PYTORCH_MPS_LOW_WATERMARK_RATIO (0.6) unless the environment already
sets them. A process-tree ceiling does not reach PyTorch's Metal allocator,
which sizes its pool against total system memory; without the watermarks the
tree is killed for an allocation the guard never had a chance to refuse.
WITH_LIMITS_MPS_HIGH_WATERMARK_RATIO and
WITH_LIMITS_MPS_LOW_WATERMARK_RATIO change the seeded ratios;
WITH_LIMITS_MPS_WATERMARKS=off disables the seeding.
WITH_LIMITS_MAX_PRESSURE, WITH_LIMITS_MIN_MEMORY_FREE_PERCENT,
WITH_LIMITS_MIN_SWAP_FREE, and WITH_LIMITS_MAX_LOAD_PER_CPU set the
admission-gate thresholds and accept the same values as their flags. The flags
take precedence over the environment.
WITH_LIMITS_RESERVATION_DIR selects the per-account reservation store. Its
default is $XDG_RUNTIME_DIR/with-limits-reservations, then
$TMPDIR/with-limits-reservations. When neither variable is set it is
/tmp/with-limits-reservations-UID on Unix, named by the effective user id
because /tmp is shared by every account, and the temporary directory on
Windows. The directory is created on demand with mode 0700, and each
supervisor writes one JSON file named by its process id. On Unix a store that
is a symbolic link, is owned by another account, or is writable by group or
others is refused: a command that would publish a reservation or wait for
headroom fails before it starts, and --check-headroom ignores the store's
records with a warning. Reservations do not coordinate across accounts, nor
across contexts of one account that resolve the store differently: on macOS a
terminal session and an ssh or launchd session see different TMPDIR
values, so set WITH_LIMITS_RESERVATION_DIR when both start guarded work.
On Windows, a Job Object provides native process-tree containment and CPU
limits. The real command is held behind a launch gate until its helper has been
assigned to the job, so it and the descendants it creates inherit containment.
On every platform, with-limits samples the resident memory of the command and
its descendants. On macOS and Linux it also samples CPU use, briefly suspending
and resuming the tracked processes to maintain the requested sustained
allowance. Wall time is supervised by with-limits on every platform.
The supervisor itself keeps its original priority so it can continue enforcing limits promptly. Only the launched command and its descendants receive the lower priority.
Sampled resident memory is the sum of resident set size (RSS) reported for the
processes in the tree. It can count shared pages more than once, so leave
headroom when processes share large mappings. Sampling also means a very brief
spike can occur between observations. Percentage and auto memory limits
additionally enforce the unallocated share as a host-memory reserve, so
unrelated or concurrently guarded growth can stop the command before its own
RSS reaches its ceiling.
On Unix the command runs in a process group of its own. The tracked tree is
every process reached from the command through parent links, plus every
member of that group, so a descendant whose parent exited between two polls
is still found as long as it stayed in the group; one that called setsid or
setpgid before it was observed is not. When the command exits while tracked
descendants remain, with-limits keeps supervising them under the same limits
and returns the command's status once the tree is empty.
A separate process group is not the terminal's foreground group, so a command
that reads from the terminal is stopped by SIGTTIN and waits there until a
time limit ends it. Background jobs, agent hooks, and CI never read the
terminal and are unaffected. For a command that must, --foreground leaves it
in the supervisor's group, as GNU timeout --foreground does: terminal input
and terminal-generated signals reach it directly, memory and time limits still
apply to the tracked tree, but a descendant that leaves the tree before it is
observed escapes containment, and a signal the terminal delivers to the group
is forwarded once more by the supervisor.
On Unix, terminating signals SIGHUP, SIGINT, SIGQUIT, and SIGTERM sent
to with-limits are forwarded to the command's process group and to tracked
descendants that have created another process group or session, followed by
SIGCONT so a stopped process receives them. with-limits then waits for
--kill-after and forcibly terminates any tracked process that remains.
SIGUSR1, SIGUSR2, and SIGWINCH are forwarded without starting
termination. SIGTSTP stops the command tree and supervisor; SIGCONT resumes
and is forwarded to the tree. Unix limit violations request graceful
termination the same way and use the same grace period. A Windows Job Object terminates the
tree as a unit. If required monitoring or enforcement fails, with-limits
stops the workload and exits with status 125.
The command's exit status is preserved unless the supervisor itself determines the result:
| Status | Meaning |
|---|---|
75 |
headroom was refused or the wait for it expired |
124 |
wall-clock limit expired |
125 |
with-limits could not supervise the command |
126 |
command was found but could not be invoked |
127 |
command was not found |
137 |
memory limit was exceeded and the process tree was terminated |
These follow GNU timeout conventions where they overlap.
- GNU
timeout(gtimeoutunder Homebrew) is mature and preferable for a time limit alone.with-limits --timefollows its duration suffixes and status124, while adding portable process-tree memory and sustained-CPU controls. - Shell
ulimitand Linuxprlimitset kernel resource limits. Their address-space and CPU-time limits are not the same as aggregate resident memory and sustained CPU rate for a process tree. - GNU
nicechanges scheduling priority on Unix.with-limitsapplies the same basic policy automatically to the whole contained tree, adds a Windows equivalent, and also enforces resource ceilings. - Linux
tasksetrestricts CPU affinity but permits full use of the selected cores.cpulimituses sampled stop/resume control similar to the Unix CPU strategy here, but is not a GNU utility and focuses on one resource. - Linux
systemd-runcan create a transient cgroup with strong native controls. It is a good Linux-specific choice when systemd and suitable delegation are available. with-gpuselects and monitors a GPU for a command. It can be used alongsidewith-limitswhen a workload needs both GPU selection and host resource containment.
Report bugs and ask usage questions in GitHub Issues. Changes can be proposed with a pull request.
with-limits is released under the MIT License.