Skip to content

Add a patch-safe file update tool for large files #3231

Description

@jasonobrown

Problem

create_or_update_file and push_files require the MCP client/model to send the complete replacement file content. For large files this can exceed host/request limits before the GitHub API receives the intended bytes.

A real private-repository workflow hit this with a ~303 KB changelog: a one-entry append required resending the entire file, the MCP request was truncated upstream of GitHub, and the resulting commit contained only ~70 KB. The bad PR was detected and closed, but a local Git client was required to recover the exact patch.

Proposed capability

Add a narrowly scoped patch-safe update tool, e.g. apply_file_patch or create_commit_from_patch, that fetches the current file server-side and applies a bounded exact patch without requiring the MCP client to transmit the entire replacement file.

A conservative initial contract could use exact text replacements rather than fuzzy patching:

  • owner, repo, branch, path
  • expected_head_sha
  • expected_blob_sha
  • ordered edits containing exact old_text and new_text
  • each edit must match exactly once (or an explicitly supplied exact occurrence count)
  • commit_message

Required behavior:

  1. Read current branch/head and fail if expected_head_sha differs.
  2. Fetch the current blob/file server-side and fail if expected_blob_sha differs.
  3. Apply edits exactly; no fuzzy offsets or best-effort matching.
  4. Preserve file mode/path and reject binary/oversized/ambiguous inputs for the initial implementation.
  5. Commit only after all edits and validations pass.
  6. Return before/after blob SHA, commit SHA, tree SHA, changed path, and size.
  7. Never force-update a ref; concurrent branch movement must fail closed.

A later version could support a bounded unified-diff parser and atomic multi-file commits through the Git database APIs.

Why this belongs in the server

The server already has the authenticated GitHub client and can fetch the full current blob directly. Keeping the full file server-side avoids model/context/request amplification and materially reduces the risk of truncation while retaining optimistic concurrency checks.

Scope / non-goals

  • No arbitrary shell or local Git execution.
  • No fuzzy patch application.
  • No force push.
  • No bypass of repository permissions or branch protections.
  • Keep the tool outside read-only mode and subject to the same OAuth/repository scope filtering as existing content-write tools.

I can prepare a focused PR with tests if this shape is acceptable.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions