The Review Methodology
Product quality and use-case fit answer different questions. An editorial review asks how well the actual product works. A Fit Score asks how well it suits a named job. Neither substitutes for the other.
Quality is not Fit Score
A product can be well made and still be the wrong choice for your budget, workflow, or task. Our review rating is an editorial judgment about product quality within an explicitly stated review scope. The separate Fit Score methodology keeps Capability, Reliability, Value, Ease, and Compatibility tied to each named, already-published use case.
A product-wide quality judgment needs representative coverage of the core product. A narrow test is not promoted into a universal verdict. Catalog-wide Fit Score coverage does not mean catalog-wide product testing or review coverage.
One shared quality framework
Every product category uses these four dimensions. They guide what we test and explain; they are not four numeric subscores to average.
Core delivery and correctness
Does the actual product do its core job correctly?
We assess the behavior and outputs the product itself is responsible for, against explicit expected results. A successful demo or a vendor feature list alone does not establish quality.
Dependability and recovery
Can the workflow be repeated, and what happens when it fails?
We look for repeatability, clear errors, safe recovery, and useful behavior after a failed attempt. Failures and retries remain part of the evidence rather than disappearing behind the best result.
Control and completeness
Can users direct, inspect, and finish the work?
We examine the controls, configuration, visibility, and input/output paths needed to complete representative tasks. Missing essentials and practical workarounds matter more than the length of a feature list.
Coherence and usability
Does the product make sense as a working whole?
We examine setup, terminology, defaults, documentation, feedback, and the path from starting a task to a usable result. Usability includes command-line and API workflows, not just visual polish.
Representative tests, specific to the category
Before testing, define the product boundary, core tasks, fixtures, expected outcomes, and failure or recovery cases. The task set must represent the product's core responsibilities, not just the easiest path to a favorable result. These are examples of test-plan design, not claims that these tests have already been performed:
- Local runtimes and APIs: installation, model or resource acquisition, request handling, documented controls, invalid requests, and recovery after failure.
- Automation tools: trigger-to-action execution, data mapping, inspection of run history, duplicate events, failed steps, and retry behavior.
- Research and writing products: source fidelity, task correctness, editing controls, and a usable export, using fixtures with checkable facts and expected results.
- Image, audio, and video tools: instruction adherence, artifact quality, editing and revision controls, and completion of a usable export in the tested formats.
Model quality is not runtime quality. A local runtime may correctly serve a model whose answers are wrong; a strong model does not prove that the surrounding application is dependable. Record and attribute both behaviors to the responsible layer where the evidence allows. If the cause is uncertain, say so rather than inventing a diagnosis.
One whole-number editorial judgment, 1 to 5
The rating synthesizes the evidence across the four dimensions. It is not a weighted formula, a Fit Score average, or a conversion from a third-party rating. A serious core failure can outweigh several polished secondary features. The review explains why the selected rating is justified, why it is not higher, and why it is not lower.
| Rating | Judgment | Evidence required |
|---|---|---|
| 5/5 | Strong across the core product | Broad, strong evidence across representative tasks supports the judgment, with no material weaknesses in the assessed core behavior. |
| 4/5 | Dependable | The assessed behavior works dependably, with bounded limitations that do not undermine its core delivery. Coverage must still be stated. |
| 3/5 | Material compromises | The product is useful, but meaningful compromises or recurring workarounds affect the assessed workflow. |
| 2/5 | Major flaws | Major flaws undermine important parts of the core job, even if some tasks can be completed. |
| 1/5 | Repeated core failure | Repeated failures prevent the product from reliably completing its core job in the tested scope. |
Unreviewed is not 0/5. Lack of access, time, or sufficient evidence is a coverage limitation, not evidence of poor product quality. Popularity and affiliate relationships do not factor into quality ratings. Price alone does not establish quality or automatically raise or lower a rating. Price and plan suitability remain relevant context and are part of the separate use-case Fit rubric.
What the evidence record must contain
- A test plan with the product scope, representative tasks, fixtures or inputs, expected outcomes, and criteria for assessing them.
- The product version or build, tested plan, configuration, environment, and relevant model, hardware, or integration versions.
- Actual observations and evidence references: outputs, logs, or captures, including failures, timeouts, retries, and any changes made between attempts.
- The review author, testing method, test date, review date, and methodology version, with enough detail to distinguish observations from interpretation.
- Limitations and untested areas, including features, workloads, platforms, plans, or environments outside the evidence.
- Rating reasons tied to the evidence, including why the judgment is not higher or lower and any unresolved uncertainty.
If agents operate the tests or review the evidence, that involvement must be disclosed. Automated or agent-operated testing is not described as human hands-on testing. We do not invent a human reviewer, test run, measurement, failure explanation, or result to fill an evidence gap.
A review is not a security, privacy, compliance, or production-readiness certification. Testing one workflow does not establish an uptime guarantee, network isolation, or reliability at an untested scale.
Review status and coverage
- Not reviewed
- No qualifying original review is available. This is not a zero or a negative judgment.
- Limited scope
- Original evidence covers only a narrow workflow or environment. Any published rating applies only to that stated scope, not the whole product.
- Reviewed
- The review has sufficient representative evidence for its declared product scope. This does not mean every feature, model, plan, or environment was tested.
- Needs update
- Product changes or aging evidence require a new assessment. Earlier conclusions are not silently treated as current.
- Blocked
- Access, setup, evidence, or another prerequisite prevents a defensible assessment. We disclose the blocker rather than manufacture a score.
The existing Ollama review is limited scope
The existing Ollama 4/5 review covers one local CLI/API workflow, with one small model on one Linux host. It is not a product-wide quality rating and does not cover every model, hardware configuration, workload, or Ollama Cloud. This methodology does not rescore that review or create scores for other products.
Material product changes, new failure evidence, or outdated test conditions require reassessment. Broader coverage requires new evidence, not a relabeling of the old review. Status, scope, and dates must stay visible so an earlier result is not mistaken for a current, universal judgment.
Publication and search presentation
Our rated-product publication rule pairs a supported offer with an original editorial review of the same product. The offer must have evidence for the actual plan, price, billing unit, source, and check date; the review must have its own evidence, author, dates, scope, and rating reasons. We do not invent a free price, a review, or a score to complete that pair.
We do not aggregate third-party reviews into our editorial rating or convert Fit Scores into review ratings. Structured data must reflect the supported offer and original review visible on the page, including the review's limits. A single editorial review is not a crowd-sourced aggregate rating.
Valid Google structured data is not a guarantee of search stars, rich results, indexing, or placement. Search presentation is controlled by the search engine, not by our rating or markup.
Choosing a tool for a particular task? Read the complementary Fit Score methodology for the five use-case axes and their one-decimal average.
Review methodology version 1.0.