-
Notifications
You must be signed in to change notification settings - Fork 16
Add article: What I Learned Building an AI Code Reviewer with Java and Spring Boot #102
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Open
Sweety717
wants to merge
11
commits into
foojayio:main
Choose a base branch
from
Sweety717:main
base: main
Could not load branches
Branch not found: {{ refName }}
Loading
Could not load tags
Nothing to show
Loading
Are you sure you want to change the base?
Some commits from the old base branch may be removed from the timeline,
and old review comments may become outdated.
Open
Changes from all commits
Commits
Show all changes
11 commits
Select commit
Hold shift + click to select a range
1902d18
Create index.md
Sweety717 0e9c365
Create index.md
Sweety717 dc9b742
Add files via upload
Sweety717 2e2da49
Delete codeguard-ai-cover.jpg.png
Sweety717 9fa6f5d
Add files via upload
Sweety717 ed18c42
Update index.md
Sweety717 0a53c6e
Update index.md
Sweety717 1cb5694
Rename codeguard-ai-cover.jpg.png to codeguard-ai-cover.png
Sweety717 7e6a862
Update and rename index.md to _index.md
Sweety717 f6fb6a9
Rename _index.md to index.md
Sweety717 6ccf60e
Rename index.md to _index.md
Sweety717 File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,6 @@ | ||
| --- | ||
| title: "IsabiTech" | ||
| description: "Java, Spring Boot and AI developer tools built for practical software engineering workflows." | ||
| --- | ||
|
|
||
| IsabiTech builds practical developer tools using Java, Spring Boot, GitHub APIs, and AI technologies. |
Binary file added
BIN
+1.24 MB
...arned-building-an-ai-code-reviewer-with-java-spring-boot/codeguard-ai-cover.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
303 changes: 303 additions & 0 deletions
303
draft/what-i-learned-building-an-ai-code-reviewer-with-java-spring-boot/index.md
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,303 @@ | ||
| --- | ||
| title: "What I Learned Building an AI Code Reviewer with Java and Spring Boot" | ||
| date: "2026-01-01" | ||
| description: "What I learned building an AI code-review workflow around GitHub Pull Requests, Spring Boot, and multiple AI providers." | ||
| authors: | ||
| - "isabitech" | ||
| image: "codeguard-ai-cover.png" | ||
| categories: | ||
| - "Java" | ||
| - "Spring" | ||
| - "AI" | ||
| - "Developer Tools" | ||
| - "GitHub" | ||
| --- | ||
|
|
||
| AI-assisted coding has changed how quickly developers can write software. | ||
|
|
||
| But that creates another question: | ||
|
|
||
| **How do you review all that generated code effectively?** | ||
|
|
||
| That question led me to build CodeGuard AI, a self-hosted AI code-review application for GitHub Pull Requests using Java and Spring Boot. | ||
|
|
||
| At first, the architecture seemed simple: | ||
|
|
||
| ```text | ||
|
coderabbitai[bot] marked this conversation as resolved.
|
||
| GitHub Pull Request | ||
| ↓ | ||
| LLM | ||
| ↓ | ||
| Review | ||
| ``` | ||
|
|
||
| In practice, that wasn't enough. | ||
| The interesting engineering work was everything around the model: collecting useful Pull Request context, defining review instructions, structuring the model response, assessing findings, and integrating the result back into the GitHub workflow. | ||
| This article explains what I learned while building that workflow. | ||
| The real problem isn't calling an LLM | ||
| Calling an AI provider from a Spring Boot application is relatively straightforward. | ||
| The harder question is: | ||
| What exactly should the model receive, and what should the application do with the response? | ||
| A Pull Request contains much more useful information than a single block of changed code. | ||
| A review workflow can involve: | ||
| - Pull Request metadata | ||
| - Changed files | ||
| - Code diffs | ||
| - Review instructions | ||
| - Repository information | ||
| - Analysis rules | ||
| - The selected AI provider | ||
| - The model's response | ||
| The workflow therefore becomes: | ||
| GitHub PR | ||
| ↓ | ||
| Fetch PR context | ||
| ↓ | ||
| Build review context | ||
| ↓ | ||
| Apply review instructions | ||
| ↓ | ||
| AI provider | ||
| ↓ | ||
| Parse response | ||
| ↓ | ||
| Structured findings | ||
| ↓ | ||
| Risk / severity / confidence | ||
| ↓ | ||
| GitHub review | ||
|
|
||
| The model is only one part of the system. | ||
| Building the GitHub integration | ||
| The first major component is the GitHub integration. | ||
| The application needs to retrieve the Pull Request information and the changes that need to be reviewed. | ||
| Conceptually, the workflow looks like: | ||
| Repository | ||
| ↓ | ||
| Pull Request | ||
| ↓ | ||
| Changed files | ||
| ↓ | ||
| Diffs | ||
| ↓ | ||
| Review context | ||
|
|
||
| The important design decision here is that the reviewer should not blindly send an entire repository to the model. | ||
| Instead, the application can start with the changes that are actually part of the Pull Request and construct the context around those changes. | ||
| That makes the review more focused and avoids treating the entire codebase as equally relevant to every review. | ||
| Turning a diff into review context | ||
| A raw diff isn't necessarily enough for useful AI feedback. | ||
| The application also needs instructions that tell the model what kind of review it is expected to perform. | ||
| For example, the review instructions can ask the model to look for: | ||
| - Bugs and logic problems | ||
| - Security issues | ||
| - Performance problems | ||
| - Best-practice violations | ||
| The important distinction is between sending code to an LLM and building a review context for an LLM. | ||
| The second approach gives the application more control over how the review is performed. | ||
| Supporting multiple AI providers | ||
| One of the design goals of CodeGuard AI was avoiding a workflow that depended on a single AI provider. | ||
| The application supports: | ||
| - OpenAI | ||
| - Google Gemini | ||
| - Ollama | ||
| That creates an abstraction between the review workflow and the provider. | ||
| Conceptually: | ||
| ┌── OpenAI | ||
| │ | ||
| Review workflow ─┼── Gemini | ||
| │ | ||
| └── Ollama | ||
|
|
||
| The application can therefore keep the GitHub and review workflow relatively independent from the model provider. | ||
| This also makes experimentation easier. | ||
| Different providers can produce different results, and a self-hosted option such as Ollama can be useful when running the model locally is important. | ||
| Raw AI output isn't enough | ||
| This was one of the biggest lessons from the project. | ||
| An AI model can return a long textual response explaining several possible problems. | ||
| But a developer reviewing a Pull Request doesn't necessarily want another large block of text. | ||
| They need to know: | ||
| - What is the finding? | ||
| - Where is it? | ||
| - How serious is it? | ||
| - How confident is the analysis? | ||
| - What could be done about it? | ||
| So the review workflow converts the model response into structured findings. | ||
| A simplified finding can be thought of as: | ||
| Finding | ||
| ├── Type | ||
| ├── Severity | ||
| ├── Confidence | ||
| ├── Location | ||
| ├── Explanation | ||
| └── Suggested fix | ||
|
|
||
| This changes the role of the AI response. | ||
| Instead of being the final product, it becomes input to a developer-oriented review workflow. | ||
| Why severity and risk matter | ||
| An AI model might identify ten possible issues. | ||
| That doesn't mean all ten deserve the same attention. | ||
| A useful review system therefore needs some way of distinguishing findings. | ||
| For example: | ||
| CRITICAL | ||
| HIGH | ||
| MEDIUM | ||
| LOW | ||
|
|
||
| Other information can also help: | ||
| Severity | ||
| Risk | ||
| Confidence | ||
| Finding type | ||
| Code location | ||
|
|
||
| The goal isn't to let the AI make the final decision. | ||
| The goal is to make the developer's review more focused. | ||
| The workflow becomes: | ||
| 10 AI findings | ||
| ↓ | ||
| Structured findings | ||
| ↓ | ||
| Risk / severity / confidence | ||
| ↓ | ||
| Prioritized review | ||
| ↓ | ||
| Human decision | ||
|
|
||
| That distinction is important. | ||
| AI first pass does not mean AI final decision. | ||
| Sending the results back to GitHub | ||
| The review becomes much more useful when it fits into the developer's existing workflow. | ||
| Instead of requiring developers to open a separate application and manually copy the findings into GitHub, CodeGuard AI can post review results back to the Pull Request. | ||
| The resulting workflow looks like: | ||
| Developer opens PR | ||
| ↓ | ||
| CodeGuard analyzes changes | ||
| ↓ | ||
| AI generates findings | ||
| ↓ | ||
| Findings are structured | ||
| ↓ | ||
| Review is posted to GitHub | ||
| ↓ | ||
| Developer evaluates the findings | ||
|
|
||
| This makes GitHub the place where the review can be consumed. | ||
| The separate dashboard can then provide additional information such as review history and trends. | ||
| The Spring Boot side | ||
| Spring Boot provides the application layer that connects these pieces. | ||
| At a high level, the application contains responsibilities for: | ||
| GitHub integration | ||
| ↓ | ||
| Review orchestration | ||
| ↓ | ||
| AI provider integration | ||
| ↓ | ||
| Finding processing | ||
| ↓ | ||
| Persistence | ||
| ↓ | ||
| Dashboard / APIs | ||
|
|
||
| Spring Security handles authentication and authorization concerns, while Spring Data/JPA provides the persistence layer. | ||
| Maven manages the Java project dependencies and build lifecycle. | ||
| The result is not just an AI API call. It is a complete application workflow around that API call. | ||
| Where the application becomes interesting | ||
| The most interesting part of the project wasn't choosing which LLM to call. | ||
| It was deciding what happens before and after the model call. | ||
| Before the model: | ||
| GitHub | ||
| ↓ | ||
| PR changes | ||
| ↓ | ||
| Review context | ||
| ↓ | ||
| Instructions | ||
|
|
||
| After the model: | ||
| AI response | ||
| ↓ | ||
| Structured findings | ||
| ↓ | ||
| Severity / risk / confidence | ||
| ↓ | ||
| GitHub review | ||
| ↓ | ||
| Developer decision | ||
|
|
||
| That surrounding workflow is where most of the application engineering lives. | ||
| What AI code review still can't solve | ||
| Building the reviewer also made the limitations clearer. | ||
| An AI reviewer can identify many useful issues, but it doesn't automatically understand every business rule behind a system. | ||
| For example, a Pull Request might be technically valid while violating an organization-specific business requirement that isn't represented in the available context. | ||
| That means AI review should be treated as a first-pass analysis, not an unquestionable authority. | ||
| Human review still matters, especially for: | ||
| - Business-critical logic | ||
| - Security-sensitive changes | ||
| - Financial workflows | ||
| - Architectural decisions | ||
| - Organization-specific rules | ||
| The quality of an AI review therefore depends not only on the model, but also on the context and instructions provided to it. | ||
| Self-hosting changes the design | ||
| Another important design consideration was deployment. | ||
| CodeGuard AI is designed as a self-hosted application rather than a hosted code-review SaaS. | ||
| That means developers can run the application on their own infrastructure and customize the source code. | ||
| The architecture can therefore look like: | ||
| GitHub | ||
| ↓ | ||
| Self-hosted CodeGuard | ||
| ↓ | ||
| AI provider | ||
| ↓ | ||
| Review results | ||
| ↓ | ||
| GitHub | ||
|
|
||
| With Ollama, the model can also be run locally: | ||
| GitHub | ||
| ↓ | ||
| CodeGuard | ||
| ↓ | ||
| Ollama | ||
| ↓ | ||
| Local model | ||
| ↓ | ||
| Review findings | ||
|
|
||
| This doesn't automatically solve every security or privacy concern, but it gives developers more control over where the application and AI workflow run. | ||
| What I would change if I built it again | ||
| The project also changed how I think about AI developer tools. | ||
| If I were starting again, I would spend even more time on the boundary between AI output and application logic. | ||
| The model will change. | ||
| The provider will change. | ||
| The prompts will change. | ||
| But the application still needs a stable workflow for turning uncertain model output into something a developer can inspect and act on. | ||
| That suggests an architecture where the AI provider is replaceable, while the review workflow remains relatively stable. | ||
| The main lesson | ||
| The biggest lesson from building CodeGuard AI was simple: | ||
| AI code review isn't just PR → LLM → review. | ||
|
|
||
| The useful system is closer to: | ||
| Pull Request | ||
| ↓ | ||
| Context | ||
| ↓ | ||
| Review instructions | ||
| ↓ | ||
| AI provider | ||
| ↓ | ||
| Structured findings | ||
| ↓ | ||
| Risk / severity / confidence | ||
| ↓ | ||
| GitHub workflow | ||
| ↓ | ||
| Human decision | ||
|
|
||
| The LLM is important, but the engineering around it is what turns an AI response into a developer tool. | ||
| That's what I found most interesting about building an AI code reviewer with Java and Spring Boot. | ||
| About CodeGuard AI | ||
| CodeGuard AI is a self-hosted AI code-review application built with Java 17+ and Spring Boot. | ||
| It supports OpenAI, Gemini, and Ollama and can analyze GitHub Pull Requests for bugs, security issues, performance problems, and best-practice violations. | ||
| The complete source code is available for developers who want to run, customize, and extend the application. | ||
| [See CodeGuard AI on Gumroad](https://javacoder716.gumroad.com/l/codeguard-ai) | ||
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win
Make the
isabitechauthor slug resolvable.The frontmatter validator rejects this slug as unknown. It requires an author profile at
content/authors/isabitech/_index.md. Add or move the profile to that path so the PR validation passes.🤖 Prompt for AI Agents
Source: Pipeline failures