LLM Output Diff Tool
Compare two to four LLM outputs side by side, diff any pair by line or by word, and compare response length in characters and words.
LLM Output Diff Tool
Recursion is a programming technique where a function calls itself to solve a problem by breaking it down into smaller, identical subproblems. Each recursive call works on a simpler version of the original problem until reaching a "base case" that can be solved directly. Here's a simple example — calculating factorial: ```python def factorial(n): if n <= 1: # Base case return 1 return n * factorial(n - 1) # Recursive case ``` When factorial(4) is called, it computes 4 * factorial(3), which computes 3 * factorial(2), and so on until reaching factorial(1) = 1. The results then "unwind" back up: 1 * 2 * 3 * 4 = 24. Recursion is preferable to iteration when: - The problem has a naturally recursive structure (trees, graphs, nested data) - You need to explore multiple branching paths (backtracking algorithms) - The recursive solution is significantly clearer and more maintainable - You're working with divide-and-conquer algorithms (merge sort, quicksort) However, iteration is generally better when performance and memory are critical, since each recursive call adds a frame to the call stack.
**Recursion** is a fundamental concept in programming where a function calls itself in order to solve a problem. The idea is to break a complex problem into smaller instances of the same problem until you reach a trivial case that can be solved directly. ### Key Components: 1. **Base Case**: The condition that stops the recursion 2. **Recursive Case**: The part where the function calls itself with modified arguments ### Simple Example: Fibonacci Numbers ```python def fibonacci(n): if n <= 0: return 0 elif n == 1: return 1 else: return fibonacci(n-1) + fibonacci(n-2) ``` This function calculates the nth Fibonacci number by summing the two preceding numbers, recursing until it hits the base cases of 0 or 1. ### When to Use Recursion Over Iteration: Recursion shines in scenarios where: - **Tree/graph traversal**: Navigating hierarchical data structures - **Divide and conquer**: Problems like merge sort or binary search - **Mathematical definitions**: Problems naturally defined recursively (factorials, Fibonacci) - **Backtracking problems**: Sudoku solvers, maze navigation, permutation generation ### Caveats: - Recursion uses more memory (call stack overhead) - Risk of stack overflow for deep recursion - Some languages optimize tail recursion, but many do not In practice, many recursive solutions can be converted to iterative ones using an explicit stack, but the recursive version is often more elegant and easier to reason about.
What this tool does
To diff two LLM JSON outputs, paste each model's response into its own panel and switch between line-level and word-level highlighting to see exactly where the responses diverge. The comparison runs entirely in your browser using a Myers diff. For JSON outputs specifically, format both responses with identical indentation first, so structural differences align line by line instead of collapsing into one long changed row.
Updated . Provided as is. Check the output before you rely on it in production.
How to use LLM Output Diff Tool
- 1
Run the same prompt on multiple models
Take your prompt (exactly the same input) and run it on GPT, Claude, Gemini, Llama, or other models. Copy each response.
- 2
Paste outputs for comparison
Input each model's output into the diff tool. Use separate sections for each model so you can compare output quality and style.
- 3
Analyze differences
The tool highlights where outputs diverge: tone, structure, accuracy, length, or logic. Look for patterns in where each model excels (coding, reasoning, creativity, etc.).
- 4
Choose the best model for your use case
Based on quality comparison, decide which model to use for production. Some models are cheaper, some faster, some more accurate. Match the model to your actual requirements.
Questions and answers
What is LLM Output Diff Tool?
What is this tool useful for compared to a regular diff checker?
For AI agents: how to call this tool
Machine-readable contract, endpoints and examples. Humans can ignore this section.
Best Path For Builders
Browser workflow
Runs instantly in the browser with private local processing and copy/export-ready output.
Browser Workflow
This tool is optimized for instant in-browser execution with local data handling. Run it here and copy/export the output directly.
/llm-output-diff/
For automation planning, fetch the canonical contract at /api/tool/llm-output-diff.json.
How do I diff two LLM JSON outputs to compare model responses?
- Record the prompt once in the prompt field, so the comparison documents what both models were actually asked.
- Paste each model's raw response into its own panel and label it with the model name (Claude, GPT, Gemini, or a custom label).
- For JSON responses, pretty-print both outputs with the same indentation before diffing — a two-space-formatted object against a minified one shows everything as changed, while identically formatted objects reveal the actual field-level differences.
- Start in line mode to spot structural drift: missing keys, reordered sections, extra fields.
- Switch to word mode to inspect value-level drift inside matching lines: numbers, enum values, phrasing.
- Check the per-panel character, word, and line counts and the unique-content view to quantify how much the responses diverge, then copy the annotated comparison out.
What should I look for when comparing model responses?
Three classes of difference matter in practice. Schema drift: one model emits a field the other omits, or wraps the payload differently — this breaks parsers and is the first thing to check. Value drift: both models return the same structure but disagree on numbers, labels, or classifications — this is a correctness question, and the diff pinpoints each disagreement. Verbosity drift: one model pads the answer with prose or markdown fences around the JSON — relevant because it affects token cost and downstream parsing.
Is this a structural JSON comparison?
No — it is a text diff (line-level or word-level Myers diff), which is usually what you want for model comparison because it also surfaces formatting and ordering differences that a key-by-key structural comparison would normalize away. If you need order-insensitive structural equality instead, sort and format both payloads first, then diff the normalized forms.
Do my model outputs leave the browser?
No. The diff algorithm runs client-side; outputs are not uploaded or stored. Model responses often quote the prompt and internal data, so a local-only comparison is the safe default when evaluating models on real production prompts.