FF Info logo FF Info

Best Betting Sites Methodology 2026 — Our 10-Point Rubric

Our best betting sites methodology exists so ratings can be audited. The rubric is public, the bands are fixed, and the testing rota is documented.

Read the ten dimensions, learn the weightings, and use the framework to challenge any review — including ours.

18+ T&Cs apply. Please gamble responsibly.

Why methodology matters more than the score

Every ranking implies a ruler. The best betting sites ranking on any consumer site is only as trustworthy as the ruler behind it. Our approach starts on the home page for best betting sites and drills into weightings here, because a score without a ruler is just an opinion in a table.

Methodology is what allows readers to challenge a reviewer without shouting. If we score an operator six out of ten and a reader believes the score should be seven, they can point at the dimension, the band and the evidence. That is a conversation we welcome; it is also a conversation impossible to have when scores are derived from vibes.

A published methodology also protects the reader from stealth changes. If we quietly widen the "strong" band on payouts to include operators who used to score "adequate", the whole table shifts. By publishing the bands with dates, any shift is visible.

Diagram of the ten scoring dimensions around a central rubric card

Methodology also protects the reader from the survivorship bias built into every ranking. Operators that stopped taking new customers, moved outside the UK, or lost licences disappear from most tables silently. A methodology page can name what happened, when, and why the operator was dropped.

The other advantage of a published methodology is that it forces the editorial team to think in bands rather than in adjectives. Bands force decisions; adjectives allow drift. We would rather be forced into difficult decisions than allowed to drift into flattering ones.

The ten dimensions explained

Ten dimensions were chosen because that is roughly the ceiling for a rubric a reader can hold in their head. Fewer dimensions oversimplify. More dimensions collapse into noise. Each dimension has a single, sharp definition and a testable manifestation.

Notice what is missing. Bonus size is not a dimension because it is trivially gamed by wagering caps. Homepage aesthetics are not a dimension because they are irrelevant to consumer outcomes. Nor is the number of sports listed on the menu — depth-of-market on the fixtures you actually bet on is what matters.

Dimensions are also mutually independent by design. A high score on payouts does not lift the terms score. A weak support score does not drag the market-depth score. Independence prevents halo effects and keeps the eventual total legible.

Scoring bands and how points convert

Each dimension is graded on five bands: fail, weak, adequate, strong, exemplary. The bands map to fractional points on a nought-to-one scale per dimension. Fail scores zero. Weak scores a quarter. Adequate scores half. Strong scores three-quarters. Exemplary scores a full point. Ten dimensions therefore yield a maximum of ten points.

The bands themselves are defined by observable evidence, not by prose. For payouts, the "strong" band is defined as three cleared withdrawals across two methods inside the operator's stated windows. "Exemplary" adds a fourth withdrawal above a large threshold cleared inside the stated window. A reader can therefore look at our evidence log for an operator and either confirm or dispute the band.

We do not use sub-decimal scoring on individual dimensions. A dimension is either weak or adequate, never "adequate-plus". The total score inherits fractions from the sum of the ten dimensions, which is why you see 7.5 rather than 7.53 in our tables.

Bands are calibrated against observed practice, not against a theoretical ceiling. If no operator in the ranking currently reaches exemplary on terms readability, we do not lower the definition of exemplary to give someone a gold star. Empty bands are informative.

Weightings and the licence veto

Weightings sit inside the bands rather than as multipliers. This is a deliberate design choice. Multiplicative weightings are easy to move and easy to hide; band definitions are harder to move without visibly rewriting the rubric.

The single exception is licence & regulation, which acts as a veto. An operator without an active UKGC licence cannot be rated at all — it is out of scope. A licence in suspension pulls the operator out of the ranking table until the suspension is lifted. This is not a weighting so much as an eligibility rule; the rubric is for UKGC operators only.

DimensionFail (0)Adequate (0.5)Exemplary (1.0)
PayoutsAny missed windowAll windows met onceRepeated within a testing month across methods
Terms readability>5,000 words with defined-term stewClear numbering, mostly plain EnglishShort, cross-referenced, plain-English throughout
Support>5 min first responseUnder 2 min, correct but genericUnder 60 s, quotes clause numbers, escalates cool-off
MarketsFixed rota < competitor floorAt or above floor with fair marginAbove ceiling with tight margin on top fixtures
MobileFrequent crashes / no RG parityStable with RG accessibleCold-start < 2 s, full feature parity

Weightings, when they exist elsewhere in a rubric, should be public. Any weighting hidden from readers is invisible power for the reviewer. Our approach avoids the temptation by moving the weight into the band definitions themselves, where a reader can see it.

The veto is also the most consequential design decision. Removing a UKGC-licensed but troubled operator from the table is a bigger editorial act than lowering its score. Vetoes should be rare and well-argued.

The testing rota

Testing must be like-for-like or comparisons collapse. Our rota fixes the fixtures used to measure market depth, the events used to measure streaming and the support windows used to measure responsiveness. If bet365 is measured against a Wednesday-night Championship game and Sky Bet is measured against a Saturday-lunchtime Premier League game, comparing them is meaningless.

Checklist of methodology steps a reviewer records against every operator
  1. Sunday morning: soft-market fixture, cash-out latency test.
  2. Wednesday afternoon: mid-tier fixture, market count and margin.
  3. Friday evening: top-tier fixture, streaming stability test.
  4. Bank holiday: support-window measurement across three questions.
  5. Rolling month-end: three-run withdrawal test across two methods.

The rota is deliberately unglamorous. Rotas are useful precisely because they are boring; they force operators to be tested on the same measurable moments regardless of which brand is fashionable this quarter.

The rota is also public. Readers can look at our published fixture list for the current quarter and, if they have the time, replicate the tests themselves against an operator we did not include. Reproducibility is what turns a ranking from an opinion into a claim.

Rota drift is a common failure mode. When testing weeks slip because life happens, an operator's score can move for reasons that have nothing to do with the operator. We monitor rota adherence and mark any period where drift exceeded a week.

What evidence we log for every score

Evidence is what turns a rating into a review. For every band we assign, we log a timestamp, a screenshot (with any personally identifiable information redacted), the method used and the observed outcome. Evidence lives in a controlled folder and is retained for the life of the ranking.

The reader-facing manifestation is compact. On each operator profile we publish a "recent evidence" strip showing the three most recent measurements per dimension. Anything older is summarised as a rolling average. That gives readers freshness without drowning them in scroll.

Evidence retention also enables re-scoring. When an operator disputes a band assignment — which happens, politely, several times a year — we can walk backward through the evidence together and either concede or confirm.

Evidence retention has a practical side benefit: it makes it much harder to smuggle a score change past readers. If the score changes but the evidence does not, we are obliged to explain the discrepancy. That discipline keeps everyone honest.

Refresh cadence and re-score triggers

Refreshes are quarterly by default. That is the shortest cadence at which withdrawal tests, support tests and market measurements can be completed cleanly across the whole ranking table. Anything faster would produce noise; anything slower would let market changes pile up unrecorded.

Unscheduled re-scores are triggered by ownership change, platform migration, a public UKGC statement, or a cluster of complaints on the ADR record. In each case we announce the trigger on the operator profile page and turn the score orange until the re-score completes, so readers know the rating is provisional.

Annual review of the rubric itself sits above the quarterly refresh. Once a year we ask whether the ten dimensions still describe what UK consumers should care about. Changes to the rubric are versioned and dated so historical scores can be interpreted against the version they were assigned under.

Provisional orange markers are a small but useful device. They tell readers that the score is in flux without hiding the operator entirely. Consumers can make a personal decision about how much weight to put on a provisional rating.

A worked example of the rubric

Consider a hypothetical mid-tier UKGC operator, "OperatorX", tested for a September window. Its withdrawals cleared inside stated windows on all three runs (strong). Terms ran to 3,800 words with numbered clauses and one defined-term issue (adequate). Support responded in under 90 seconds and quoted clause numbers (strong). Markets sat at the fixture-rota floor with fair margin (adequate). Streaming covered four of five events cleanly (strong). Mobile had feature parity but cold-started slowly on the mid-range Android (adequate). Fees were transparent (strong). Complaints trend was flat with fast responses (strong). Responsible-gambling tooling was full and cross-linked (exemplary).

Adding those bands: 0.75 + 0.5 + 0.75 + 0.5 + 0.75 + 0.5 + 0.75 + 0.75 + 1 = 6.25 across nine dimensions, plus 1 for a live licence = 7.25 out of ten. That is a strong-but-not-exemplary operator, and the evidence explains why in a way a reader can challenge.

OperatorX above is deliberately unnamed because it is a composite; the numbers are chosen to illustrate the arithmetic rather than to describe any specific operator. Real profiles use the same arithmetic but with the operator named and the evidence dated.

Limits and honest disclosures

Every rubric has blind spots. Ours does not directly rate customer feel, brand history, promotional creativity or bettor culture — because those are subjective. Nor does it rate return-to-player figures published by the operator, because those are self-declared and unaudited. We prefer to under-claim than to smuggle unverifiable numbers into a scored table.

We also cannot rate what we cannot test. Some operators impose account-opening frictions that make full testing impossible; we mark those "not rated" with a short explanation rather than guessing a score. Not-rated is more honest than a manufactured number.

Blind spots are not shameful; hidden blind spots are. By publishing where the rubric does not reach, we invite readers to weigh those omissions rather than assume the rubric covers everything a reasonable consumer would care about.

Consumers should treat every ranking table as a working hypothesis rather than as a settled fact. Working hypotheses invite challenge; settled facts invite complacency. Our rubric is designed to be challenged, and every challenge that produces a better score is a small win for the reader.

The consumer-editorial contract is simple: the editorial team promises to publish evidence and to admit error; the consumer promises to read carefully and to challenge the evidence. Neither side can substitute for the other and both are needed for the ranking to be useful.

Editorial independence is not a slogan. It is a set of workflows: rankings signed off by an editor who does not handle commercial arrangements, disclosure lines dated and versioned, and a corrections page that stays public even when it is uncomfortable.

Independence sits alongside methodology as the other half of trustworthy editorial. Methodology tells the reader how the reviewer reached the score; independence tells them why the reviewer had no incentive to reach a different one.

Readers who use these pages as reference should bookmark the specific dimension pages rather than the homepage, because dimension pages are updated when definitions change and the homepage is updated on refresh cadence. Both are useful; the difference matters when you are trying to reconstruct why a score moved.

UK consumers who write to us with challenges rarely regret the exchange. Even where the challenge does not move the score, the reasoning made public in the reply is often more useful than the original ranking cell. That conversational transparency is baked into the site by design.

None of this replaces a personal risk assessment. Rankings are inputs to a decision; they do not make the decision for you. The reader retains the final call and, because the reader retains the final call, the reader also retains the responsibility for setting their own limits.

Frequently Asked Questions

How often is the rubric updated?

Rubric bands are reviewed annually. Scores against the rubric are refreshed quarterly, with unscheduled re-scores triggered by material events.

Is the methodology open source?

Yes. Adopt the ten dimensions and bands for your own reviews without permission, provided you disclose the source.

How do you avoid single-tester bias?

Testing is logged with timestamps and screenshots so a second reviewer can replicate the runs. Support-quality scores are cross-checked by a second tester on the same rota.

Do you re-score operators after a major change?

Yes. A platform migration, ownership change or public UKGC statement triggers an unscheduled re-score and a provisional orange marker on the profile.

What weight sits on the licence dimension?

Effectively a veto. A missing or suspended UKGC licence removes the operator from the ranking table until the licence is reinstated.

Do you publish raw test logs?

We publish the testing rota and summary logs. Personal-account data is redacted for security. Aggregated evidence is available to any reader who challenges a score.

Responsible Gambling

Betting is entertainment for consumers who can afford to lose. If your play is affecting sleep, work or relationships, take a break and reach out. UK support is available from GamCare, GordonMoody, the NHS gambling clinic network and BeGambleAware. Self-exclusion across UKGC-licensed operators sits with GamStop. These references are text-only from our pages by network policy. 18+ T&Cs apply. Please gamble responsibly. Background reading is available at Wikipedia and gov.uk.

Illustrated portrait of Alex Ridley

Alex Ridley — Consumer Editor

Alex has built operator-review frameworks since 2018, testing what a fair, verifiable UK betting-site review actually looks like when affiliate incentives are stripped out.