• 
      

    Improve the way we fetch and build changes from GitHub.

    Review Request #15261 — Created Oct. 9, 2026 and updated

    Information

    Review Board
    release-9.x

    Reviewers

    This change rearranges some methods and adds new ones to the GitHub client
    and hosting service to improve the way we fetch and build GitHub changes.
    The logic for all of this used to live in GitHub.get_change(), and now
    we've pulled it out into general helper methods and improved the way
    we build and process GitHub diffs.

    We have a new GitHubClient.get_diff_content() method
    that handles getting the diff content for a change and expanding the
    blob SHAs in it. get_change() used to build the diff manually by
    looping over the files in the change, but now we fetch the raw diff
    content directly from GitHub's API and replace the index lines with
    full SHA index lines, in pretty much the same way that we do for
    Forgejo.

    Our old code also didn't handle commits without a parent, aka the root
    commit of a repo, but now we handle this and allow for review requests
    to be created from these commits. And files without textual changes,
    like binary files, mode-only changes, and pure renames are now included
    in the diff, whereas before they were dropped. We still don't include
    the actual content of binary file changes, but we could add that later.

    We also now cache blob SHAs onto the client and in the cache backend,
    to help reduce the amount of API requests we make for file
    existence checks and diff processing. This will be important for pull
    request integration, where we're likely to fetch a lot of commits and
    diffs at once.

    The get_change() unit test was updated to use a raw diff from GitHub
    instead of the JSON from before, and checks for a binary file and
    a rename header.

    • Ran unit tests.
    • Used for GitHub pull request integration.
    • Tested creating a new review request from a GitHub commit through
      the UI (a root commit, and a commit that has binary file and .py
      file changes).
    Summary ID
    Improve the way we fetch and build changes from GitHub.
    This change rearranges some methods and adds new ones to the GitHub client and hosting service to improve the way we fetch and build GitHub changes. The logic for all of this used to live in `GitHub.get_change()`, and now we've pulled it out into general helper methods and improved the way we build and process GitHub diffs. We have a new `GitHubClient.get_diff_content()` method that handles getting the diff content for a change and expanding the blob SHAs in it. `get_change()` used to build the diff manually by looping over the files in the change, but now we fetch the raw diff content directly from GitHub's API and replace the index lines with full SHA index lines, in pretty much the same way that we do for Forgejo. Our old code also didn't handle commits without a parent, aka the root commit of a repo, but now we handle this and these commits can be reviewed. And files without textual changes, like binary files, mode-only changes, and pure renames are now included in the diff, whereas before they were dropped. We still don't include the actual content of binary file changes, but we could add that later. We also now cache blob SHAs onto the client and in the cache backend, to help reduce the amount of API requests we make for file existence checks and diff processing. This will be important for pull request integration, where we're likely to fetch a lot of commits and diffs at once. The `get_change()` unit test was updated to use a raw diff from GitHub instead of the JSON from before, and added checks for a binary file and a rename header.
    48ccf94187fd0965957c0d22af5aa9ab25b7c1a4
    Checks run (2 succeeded)
    flake8 passed.
    JSHint passed.