Skip to content

Latest commit

 

History

31 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

with-limits

Crates.io CI License: MIT Rust 1.85+

with-limits keeps a development machine responsive while background commands run. It gives foreground applications scheduling priority, limits the child process tree's memory, sustained CPU use, and wall-clock runtime, and works on macOS, Windows, and Linux.

The motivating case is concurrent agents doing machine-learning or research work. Several agents can each start a RAM-intensive job that looks reasonable in isolation, while their combined memory use exhausts and crashes the host. Putting a process-tree limit around each job contains that failure.

The entire child tree also runs at a fixed lower scheduling priority. On Unix, its inherited niceness is increased by 10, capped at the maximum of 19. Windows uses Below Normal priority. This still lets a job use otherwise-idle CPU capacity while allowing interactive applications to win when they need it.

For a time limit alone, GNU timeout (installed as gtimeout by Homebrew) is the established choice. with-limits is useful when one portable command should also constrain memory or sustained CPU consumption.

Status

with-limits is an early-stage, actively maintained tool. Releases before 1.0 may change command-line options, environment variables, and exit-status behavior.

Installation

Installation requires Rust 1.85 or newer and Cargo. The Rust toolchain installer provides both.

Install from crates.io:

cargo install with-limits --locked

This installs with-limits in Cargo's binary directory, normally ~/.cargo/bin. Ensure that directory is on your PATH.

Build from Source

git clone https://github.com/osteele/with-limits.git
cd with-limits
cargo install --path .

with-limits is intended for trusted local and CI workloads. It is a resource guard, not a security sandbox for hostile code.

Usage

Verify the installation with a command available alongside Cargo:

with-limits --time 5s -- cargo --version

A successful run prints the Cargo version and exits with status zero. In an interactive terminal, the startup summary also reports the time limit and lower scheduling priority.

Place the command after --:

with-limits --memory 8GiB --cpu 2 --time 30m -- python train.py

Use -c to run a shell command. On macOS this uses /bin/zsh; on other Unix systems it uses /bin/sh; and on Windows it uses %COMSPEC%, falling back to cmd.exe when that variable is unset.

with-limits --memory 4GiB -c 'just format && just check'

When no limit option is supplied, with-limits applies --memory auto, which is 70% of the memory available when the command starts. It also preserves the remaining 30% as a continuously checked host reserve. This lets several concurrent guarded agents react to their combined memory use:

with-limits -c 'just format && just check'

Lower scheduling priority is independent of the resource-limit options and is enabled for every command by default.

Options:

  • --memory SIZE, -m SIZE limits aggregate resident memory. Sizes accept SI suffixes such as GB, IEC suffixes such as GiB, auto, or a percentage of initially available memory such as 60%. On macOS, available memory is the kernel's free percentage (kern.memorystatus_level) of physical memory. A percentage or auto limit that resolves below 256 MiB exits 75 without running the command.
  • --reserve SIZE publishes an absolute memory estimate for admission, independently of the containment cap. An explicit absolute --memory size is also the reservation when this option is absent.
  • --cpu CORES limits sustained CPU use, where 1 is the capacity of one logical core. Fractional values such as 0.5 are accepted; the minimum is 0.01.
  • --time DURATION, -t DURATION limits wall-clock runtime. Durations use the GNU timeout suffixes s, m, h, and d; milliseconds use ms.
  • --kill-after DURATION (also --grace) controls how long graceful termination may take before the process tree is forcibly terminated. The default is two seconds.
  • --shell PATH selects the shell used by -c.
  • --require-native rejects a requested memory or CPU limit if the platform would enforce it by sampling. Memory limits are currently sampled on every platform; Windows CPU limits are native.
  • --foreground keeps the command in the supervisor's process group so it can read from the terminal. See Enforcement for the tradeoff. Unix only; accepted and ignored on Windows.
  • --quiet, -q suppresses the interactive startup summary. Limit violations are still reported.
  • --no-reservation ignores outstanding admission reservations and does not publish one for the command.

Admission gate

Containment limits a command after it starts. The admission gate answers the earlier question of whether to start one at all, so several independent agents can consult the same host policy instead of each keeping its own copy. The resource that exhausts first on a machine running many agents is RAM, so the gate is built on memory signals: CPU load is the symptom of paging, and a load threshold fires late and on the wrong quantity.

Four signals feed the decision, and two of them are enforced by default.

The kernel memory pressure level (1 normal, 2 warn, 4 critical) is read on macOS from kern.memorystatus_vm_pressure_level and on Linux derived from the PSI /proc/pressure/memory stall time — a mapping onto the macOS scale chosen for this tool, not a kernel verdict; it is unknown on Windows. The default refuses only critical. Warn is the ordinary operating state of a machine running many concurrent tasks: measured on one such workstation it persisted for hours at a stretch while a third of memory stayed free, so refusing warn parks work indefinitely without reducing risk.

Free memory as a percentage of total is the signal that separates ordinary memory management from exhaustion, and the default floor is 10%. On macOS it is the kernel's own kern.memorystatus_level, the figure Apple's memory_pressure tool prints as the free percentage. It counts every page that is neither wired nor held by the compressor, active pages included, so it is well above the vm_stat free count: a host with 0.4% of pages free and 40% in the compressor reported 31%. Elsewhere it is derived from available and total memory. On the workstation above, every state the gate refused sat between 27% and 53% free, while the state it exists to refuse had 32 MB free, effectively zero.

Swap free and swap total are read on every platform and are reported, not enforced, by default. Free space inside the current swap files does not measure exhaustion on macOS: swap is sized on demand and grows while free disk allows, so the fraction in use rises simply because the kernel sized swap to the workload. Setting --min-swap-free enforces a floor for a platform where swap is a fixed partition; a swap total at or below that floor is then exempt, since swap that small has barely been touched.

The one-minute load average per logical CPU is read on macOS and Linux and is unknown on Windows; it is reported but not enforced unless a ceiling is set. CPU load is the late symptom of paging, not the resource that runs out.

Every signal is optional. A platform that cannot answer, a failed sysctl, or an unparsable file yields an unknown for that signal only, and unknown is never replaced by a fabricated value. An unknown signal is reported by name but does not refuse, because whether "cannot tell" is a refusal is the caller's policy; --refuse-unknown treats an unknown enforced signal as a refusal for the caller that wants that.

An explicit absolute reservation publishes an admission promise. For example, --memory auto --reserve 3GiB allows the tree to use the automatic containment cap while asking admission checks to assume about 3 GiB of growth. Other admission checks subtract the unrealized part of that promise from free memory. The unrealized amount is the reservation minus the tree's most recently sampled RSS, with a floor of zero. A tree that grows past its reservation contributes no unrealized bytes but remains governed by its independent memory cap. An observation older than 30 seconds contributes the full reservation.

The supervisor refreshes and retains its reservation while the store remains available. A refresh never waits for the store lock: when another process holds it, that tick's refresh is skipped. A refresh failure prints one warning, removes the reservation, and continues supervising and terminating the command. Each record includes the supervisor's process start time so a recycled process id does not keep an orphaned reservation alive. Platforms that cannot read a start time fall back to process-id liveness.

The admission decision and reservation write share one exclusive file lock. Two callers that reach the gate together therefore make their decisions in sequence, and the second sees the first caller's promise. The lock is held only while reading, deciding, and writing. It is released before the command or a headroom wait begins.

auto and percentage memory limits do not implicitly publish reservations. They cap a tree relative to whatever memory is free, so using the cap as the promise would let the first caller claim most of the host and serialize later work. Pair either form with an absolute estimate when admission coordination is required:

with-limits --wait-for-headroom=15m --memory auto --reserve 3GiB -- command

An absolute --memory size remains an implicit reservation when --reserve is absent. --reserve is valid without --memory. When it is the only resource option, the usual default --memory auto containment applies. If --cpu or --time is also present, omitting --memory leaves memory uncapped. --check-headroom reads reservations but never creates one.

with-limits --check-headroom reads the signals and outstanding reservations, applies the policy, prints one line per signal plus the reservation count and unrealized byte total, and exits 0 when admitted or 75 (EX_TEMPFAIL) when refused. --json prints the readings, the effective policy, the reservation summary, the refusing reasons, and the unknown signals as one JSON object instead:

with-limits --check-headroom
with-limits --check-headroom --json

with-limits --wait-for-headroom[=DURATION] [limits] -- command polls the same decision until admitted, then runs the command under the given limits. The wait is unbounded without a DURATION; with one, expiry exits 75 without running the command. Polls back off from 15 seconds by a factor of 1.5 with a random multiplier in [0.5, 1.5), capped at 90 seconds per sleep. The jitter matters because dozens of agent sessions poll independently: a fixed backoff makes them retry in lockstep and start together the moment pressure clears. While waiting, a line naming the refusing signals goes to stderr on the first refusal and at most once a minute, unless --quiet. A terminating signal during the wait ends it, since no command is running yet.

The wait never implies a limit and a limit never implies a wait; with no limit options the default --memory auto applies as usual.

  • --max-pressure N sets the highest admitted pressure level. The default is 2, so only critical pressure refuses.
  • --min-memory-free-percent PERCENT sets the free-memory floor. The default is 10; 0 reports free memory without enforcing it.
  • --min-swap-free SIZE sets the swap floor as an absolute size (percentages are not accepted). The default is 0, which reports swap without enforcing it.
  • --max-load-per-cpu CORES enforces a load ceiling. Without it, load is reported but never refuses.
  • --refuse-unknown treats an unknown enforced signal as a refusal.
  • --reserve SIZE sets the absolute estimate published when the command is admitted. It overrides the implicit reservation from an absolute --memory size and is valid without an explicit memory cap.
  • --no-reservation bypasses the reservation store for this invocation. It neither subtracts other reservations nor publishes its own.

Agent hooks

Agent lifecycle hooks can place selected shell workloads under with-limits without changing the command the model generates. The included example wraps POSIX uv run requests, including the rest of a pipeline or compound command, under one default memory budget:

uv run python experiment.py && just process-results
→ with-limits -c 'uv run python experiment.py && just process-results'
Agent Automatic wrapping path
Claude Code Native PreToolUse command hook
Codex Native PreToolUse command hook
OpenCode tool.execute.before plugin
Kimi Code CLI Session-wide launcher or command-specific PATH wrapper

Claude Code and Codex can use the same hook script. OpenCode reaches it through a small plugin. Kimi hooks cannot currently replace a tool input, so they cannot apply this command-by-command rewrite. See Agent shell hooks for the example, configuration, security tradeoff, and platform scope.

Environment

WITH_LIMITS_NICE controls fixed priority lowering. It defaults to 10. On Unix, values from 1 through 19 are added to the child process's inherited niceness, capped at 19; descendants inherit the result. On Windows, any enabled value selects the Below Normal priority class for the Job Object and therefore its entire process tree. Set it to 0, off, false, or no to disable priority lowering.

The priority does not change in response to load or the number of running agents. Fixed lower priority lets foreground work preempt background jobs while leaving idle CPU capacity available to them.

On macOS, guarded commands receive PYTORCH_MPS_HIGH_WATERMARK_RATIO (0.7) and PYTORCH_MPS_LOW_WATERMARK_RATIO (0.6) unless the environment already sets them. A process-tree ceiling does not reach PyTorch's Metal allocator, which sizes its pool against total system memory; without the watermarks the tree is killed for an allocation the guard never had a chance to refuse. WITH_LIMITS_MPS_HIGH_WATERMARK_RATIO and WITH_LIMITS_MPS_LOW_WATERMARK_RATIO change the seeded ratios; WITH_LIMITS_MPS_WATERMARKS=off disables the seeding.

WITH_LIMITS_MAX_PRESSURE, WITH_LIMITS_MIN_MEMORY_FREE_PERCENT, WITH_LIMITS_MIN_SWAP_FREE, and WITH_LIMITS_MAX_LOAD_PER_CPU set the admission-gate thresholds and accept the same values as their flags. The flags take precedence over the environment.

WITH_LIMITS_RESERVATION_DIR selects the per-account reservation store. Its default is $XDG_RUNTIME_DIR/with-limits-reservations, then $TMPDIR/with-limits-reservations. When neither variable is set it is /tmp/with-limits-reservations-UID on Unix, named by the effective user id because /tmp is shared by every account, and the temporary directory on Windows. The directory is created on demand with mode 0700, and each supervisor writes one JSON file named by its process id. On Unix a store that is a symbolic link, is owned by another account, or is writable by group or others is refused: a command that would publish a reservation or wait for headroom fails before it starts, and --check-headroom ignores the store's records with a warning. Reservations do not coordinate across accounts, nor across contexts of one account that resolve the store differently: on macOS a terminal session and an ssh or launchd session see different TMPDIR values, so set WITH_LIMITS_RESERVATION_DIR when both start guarded work.

Enforcement

On Windows, a Job Object provides native process-tree containment and CPU limits. The real command is held behind a launch gate until its helper has been assigned to the job, so it and the descendants it creates inherit containment. On every platform, with-limits samples the resident memory of the command and its descendants. On macOS and Linux it also samples CPU use, briefly suspending and resuming the tracked processes to maintain the requested sustained allowance. Wall time is supervised by with-limits on every platform.

The supervisor itself keeps its original priority so it can continue enforcing limits promptly. Only the launched command and its descendants receive the lower priority.

Sampled resident memory is the sum of resident set size (RSS) reported for the processes in the tree. It can count shared pages more than once, so leave headroom when processes share large mappings. Sampling also means a very brief spike can occur between observations. Percentage and auto memory limits additionally enforce the unallocated share as a host-memory reserve, so unrelated or concurrently guarded growth can stop the command before its own RSS reaches its ceiling.

On Unix the command runs in a process group of its own. The tracked tree is every process reached from the command through parent links, plus every member of that group, so a descendant whose parent exited between two polls is still found as long as it stayed in the group; one that called setsid or setpgid before it was observed is not. When the command exits while tracked descendants remain, with-limits keeps supervising them under the same limits and returns the command's status once the tree is empty.

A separate process group is not the terminal's foreground group, so a command that reads from the terminal is stopped by SIGTTIN and waits there until a time limit ends it. Background jobs, agent hooks, and CI never read the terminal and are unaffected. For a command that must, --foreground leaves it in the supervisor's group, as GNU timeout --foreground does: terminal input and terminal-generated signals reach it directly, memory and time limits still apply to the tracked tree, but a descendant that leaves the tree before it is observed escapes containment, and a signal the terminal delivers to the group is forwarded once more by the supervisor.

On Unix, terminating signals SIGHUP, SIGINT, SIGQUIT, and SIGTERM sent to with-limits are forwarded to the command's process group and to tracked descendants that have created another process group or session, followed by SIGCONT so a stopped process receives them. with-limits then waits for --kill-after and forcibly terminates any tracked process that remains. SIGUSR1, SIGUSR2, and SIGWINCH are forwarded without starting termination. SIGTSTP stops the command tree and supervisor; SIGCONT resumes and is forwarded to the tree. Unix limit violations request graceful termination the same way and use the same grace period. A Windows Job Object terminates the tree as a unit. If required monitoring or enforcement fails, with-limits stops the workload and exits with status 125.

Exit status

The command's exit status is preserved unless the supervisor itself determines the result:

Status Meaning
75 headroom was refused or the wait for it expired
124 wall-clock limit expired
125 with-limits could not supervise the command
126 command was found but could not be invoked
127 command was not found
137 memory limit was exceeded and the process tree was terminated

These follow GNU timeout conventions where they overlap.

Related tools

  • GNU timeout (gtimeout under Homebrew) is mature and preferable for a time limit alone. with-limits --time follows its duration suffixes and status 124, while adding portable process-tree memory and sustained-CPU controls.
  • Shell ulimit and Linux prlimit set kernel resource limits. Their address-space and CPU-time limits are not the same as aggregate resident memory and sustained CPU rate for a process tree.
  • GNU nice changes scheduling priority on Unix. with-limits applies the same basic policy automatically to the whole contained tree, adds a Windows equivalent, and also enforces resource ceilings.
  • Linux taskset restricts CPU affinity but permits full use of the selected cores. cpulimit uses sampled stop/resume control similar to the Unix CPU strategy here, but is not a GNU utility and focuses on one resource.
  • Linux systemd-run can create a transient cgroup with strong native controls. It is a good Linux-specific choice when systemd and suitable delegation are available.
  • with-gpu selects and monitors a GPU for a command. It can be used alongside with-limits when a workload needs both GPU selection and host resource containment.

Questions and contributions

Report bugs and ask usage questions in GitHub Issues. Changes can be proposed with a pull request.

License

with-limits is released under the MIT License.

About

Run commands with portable process-tree limits on CPU, memory, and wall-clock time

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages