SeedVR2 vs Topaz Wonder 3: A 264-Pair Blind Evaluation of Two Image Upscalers

In the research, 13 evaluators compared SeedVR2 and Topaz Wonder 3 on 264 image pairs and tied 54-60% of them. SeedVR2 leads on edge sharpness, Topaz Wonder 3 wins faces, and the free model is the defensible default.

Two blindfolded businessmen sitting opposite each other in a futuristic blue room.
Two men in dark business suits and black blindfolds sit face to face in a minimalist blue interior, creating a tense, surreal atmosphere of uncertainty and hidden information.

Evaluation period: July 2026 · Session: ff12fe31 · Confidence level: directional

DIRECTIONAL RESULT

Inter-rater agreement in this session (Gwet’s AC1 = 0.06–0.12) falls below the threshold Everypixel requires for a verified label. Every aggregate figure below describes direction, not an attested value. The reason for the low agreement is itself a finding and is reported in Section 3.5.

Abstract

SeedVR2 and Topaz Wonder 3 are statistically close to indistinguishable in blind pairwise evaluation. Across 264 image pairs and 791 scored comparisons by 13 professional evaluators, 54–60% of pairs were scored as ties. SeedVR2 holds a small but interval-separated advantage on edge sharpness (P = 0.60, 95% CI 0.57–0.62) and detail recovery (P = 0.55, CI 0.52–0.57); naturalness is not separated (P = 0.51, CI 0.48–0.54). Overall P(SeedVR2 > Topaz Wonder 3) = 0.55 (CI 0.53–0.58). One category reverses the ranking: on faces, Topaz Wonder 3 wins naturalness on 52% of pairs against 14%. (Everypixel production team, July 2026)


1. Introduction

Two upscalers are compared here: SeedVR2, a one-step diffusion restoration model from ByteDance with published weights that runs locally at no license cost, and Topaz Wonder 3, a paid commercial upscaler with a tunable parameter interface. Both are applied to still-image upscaling of production stock content.

The comparison follows the same protocol Everypixel used in June 2026 to compare SeedVR2 against FLUX.2 [klein] 9B: identical 264-pair structure, identical category split, the same 13-person evaluator panel. That earlier session separated the two models on every scale. This one did not, and the difference between the two outcomes is itself the primary result reported here.

Two consequences follow. First, evaluator agreement in this session fell to Gwet’s AC1 = 0.06–0.12, against 0.42–0.71 in the June session, which places the result below the threshold for a verified label under the internal evaluation protocol. It is published as directional. Second, because the two models are separated by less than the panel can reliably resolve on a single pair, the practical selection criteria shift to cost, category fit, and failure mode rather than aggregate quality.


2. Methods

2.1 Study design

Study design flow: 264 image pairs across six categories and two source types are upscaled by SeedVR2 and Topaz Wonder 3, rated blind in A/B/tie format by 13 evaluators in three roles, three ratings per pair on three quality scales, producing 791 comparisons, which pass through a Gwet AC1 agreement gate that fails and labels the result directional, and a Bradley-Terry aggregation giving P equals 0.55.
Figure 1. Evaluation design and confidence gate. Session ff12fe31, July 2026.

The test set comprised 264 image pairs: 132 real photographs and 132 AI-generated images, distributed across six content categories (crowds and complex scenes, faces, materials and textures, nature, text and signage, urban scenes), 44 pairs per category.

Each pair was presented blind in randomized left/right order and scored in A/B/tie format on three quality scales:

ScaleDefinition
detail_recoveryRecovery of structure present in the source but not resolved at source resolution
edge_sharpnessDefinition of edges and boundaries without visible sharpening artifacts
naturalnessAbsence of synthetic or over-processed appearance; plausibility of the result as a photograph

Thirteen evaluators participated across three roles: 4 QA testers, 4 art directors, and 5 distribution reviewers. Each pair was targeted for three independent ratings; 791 of 792 planned comparisons were completed.

2.2 Statistics

Inter-rater agreement is reported as Gwet’s AC1 rather than Fleiss κ. Ties account for roughly 60% of outcomes, and under a distribution this skewed the Fleiss κ expected-agreement term inflates, and κ collapses toward zero independently of actual rater concordance, a known artifact described as the paradox of kappa (Feinstein & Cicchetti, 1990). AC1 is stable under skew (Gwet, 2002). Fleiss κ is reported alongside AC1 for completeness and reads 0.00–0.05 in this session, which is an artifact rather than a measurement.

Aggregate model strength uses the Bradley-Terry model (Bradley & Terry, 1952), fitted per scale and overall, with 95% confidence intervals from bootstrap resampling. Reported probabilities are P(SeedVR2 preferred over Topaz Wonder 3) in a single comparison; a confidence interval containing 0.50 indicates no separation.

Win percentages in Sections 3.1 and 3.2 are pair-level (n = 264 per scale, aggregating the three ratings per pair). Percentages in Section 3.4 are vote-level since role subsets have unequal rating counts (n = 303 distribution reviewers, 244 art directors, and 244 QA testers).

3. Results

3.1 Overall

 Stacked bars of outcome distribution per scale: detail recovery 25 percent SeedVR2, 60 percent tie, 14 percent Topaz Wonder 3; edge sharpness 35, 54, 11; naturalness 21, 59, 20.
Figure 2. Outcome distribution by quality scale, pair-level (n = 264 per scale). Ties are the majority outcome on all three scales.
Scale SeedVR2 Tie Topaz W3 AC1 Fleiss κ P [95% CI]
Detail recovery 25% (67) 60% (159) 15% (38) 0.12 0.05 0.55 [0.52–0.57]
Edge sharpness 35% (93) 54% (142) 11% (29) 0.06 0.00 0.60 [0.57–0.62]
Naturalness 21% (56) 59% (156) 20% (52) 0.08 0.04 0.51 [0.48–0.54]
Overall 0.55 [0.53–0.58]

Pair-level, n = 264 per scale. P = Bradley-Terry probability that SeedVR2 is preferred in a single comparison. Percentages are allocated by largest remainder, so each row sums to 100; counts are exact, so a cell may sit up to one point off its count. (Everypixel production team, July 2026)

Pair-level, n = 264 per scale. P = Bradley-Terry probability that SeedVR2 is preferred in a single comparison. (Everypixel production team, July 2026)

Across 264 image pairs evaluated in July 2026, SeedVR2 and Topaz Wonder 3 were scored as ties on 54–60% of comparisons depending on the scale, and the overall Bradley-Terry probability that SeedVR2 is preferred in a given comparison is 0.55 (95% CI 0.53–0.58). (Everypixel production team, July 2026)
Dot and interval plot of Bradley-Terry probabilities with 95 percent confidence intervals: overall 0.55, detail recovery 0.55, edge sharpness 0.60, naturalness 0.51 with an interval that contains 0.50.
Figure 3. Bradley-Terry probability that SeedVR2 is preferred, with 95% bootstrap confidence intervals. The naturalness interval contains 0.50.
Edge sharpness is the only scale with a substantial separation between the two models: SeedVR2 was preferred on 35% of pairs against 11% for Topaz Wonder 3, giving P = 0.60 (95% CI 0.57–0.62). (Everypixel production team, July 2026)
Naturalness shows no separation: SeedVR2 was preferred on 21% of pairs and Topaz Wonder 3 on 20%, and the Bradley-Terry confidence interval (0.48–0.54) contains 0.50. (Everypixel production team, July 2026)

The overall effect is real but small. A probability of 0.55 corresponds to a preference expressed in roughly one comparison in twenty above chance against a tie rate near 60%.

3.2 By content category

Three small-multiple panels, one per quality scale, showing win-rate margin per category. Materials and textures reaches plus 64 points for SeedVR2 on edge sharpness; faces reaches minus 39 points for Topaz Wonder 3 on naturalness.
Figure 4. Win-rate margin by category and scale (SeedVR2 % − Topaz Wonder 3 %, pair-level, n = 44 per category). Blue indicates SeedVR2 preferred, orange Topaz Wonder 3 preferred.
Category Detail A/B/tie Sharpness A/B/tie Naturalness A/B/tie
crowds_complex 41 / 11 / 48 29 / 14 / 57 43 / 16 / 41
faces 18 / 25 / 57 36 / 21 / 43 14 / 52 / 34
materials_textures 25 / 7 / 68 68 / 5 / 27 16 / 11 / 73
nature 18 / 14 / 68 36 / 5 / 59 11 / 21 / 68
text_signage 20 / 16 / 64 25 / 11 / 64 18 / 14 / 68
urban_scenes 29 / 14 / 57 16 / 11 / 73 25 / 5 / 70

A = SeedVR2, B = Topaz Wonder 3. Percentages of pairs within the category (n = 44). Percentages are allocated by largest remainder, so each row sums to 100; counts are exact, so a cell may sit up to one point off its count. (Everypixel production team, July 2026)

A = SeedVR2, B = Topaz Wonder 3. Percentages of pairs within the category (n = 44). (Everypixel production team, July 2026)

Two category effects exceed the noise floor of the overall result.

SeedVR2’s edge-sharpness advantage concentrates in materials and textures, where it was preferred on 68% of pairs against 5% for Topaz Wonder 3, the largest single-cell margin in the July 2026 session. (Everypixel production team, July 2026)
Faces reverse the overall ranking: Topaz Wonder 3 was preferred on naturalness for 52% of face pairs against 14% for SeedVR2, and also led on detail recovery, 25% to 18%. (Everypixel production team, July 2026)

The face result is consistent with the qualitative reports in Section 4. Topaz Wonder 3 renders facial skin as a continuous, even surface, which raters scored as natural at portrait scale. SeedVR2 resolves pore-level structure and, when the reconstruction is incorrect, places the error on a face. The two models also trade in the opposite direction on a second face attribute: raters reported that SeedVR2 preserves facial proportions and likeness more closely, while Topaz Wonder 3 alters features more visibly. Naturalness and fidelity are not the same measurement, and on faces the two models split them.

3.3 By source type

Source type Detail A / tie Sharpness A / tie Naturalness A / tie
Real photographs (n = 132) 29% / 60% 35% / 56% 19% / 64%
AI-generated (n = 132) 22% / 61% 36% / 52% 23% / 55%

A = share of pairs on which SeedVR2 was preferred. (Everypixel production team, July 2026)

A = share of pairs on which SeedVR2 was preferred. (Everypixel production team, July 2026)

Source origin does not modify the result: SeedVR2’s edge-sharpness win rate is 35% on real photographs and 36% on AI-generated sources, and tie rates differ by no more than 6 percentage points on any scale. (Everypixel production team, July 2026)

This is a negative result worth recording. In the June 2026 session against FLUX.2 [klein] 9B, source origin identified the one condition under which the losing model won. Here it carries no decision value, and a mixed-content pipeline does not need to route by source type.

3.4 Between-role variation

Two panels of vote distribution by evaluator role. On edge sharpness all three roles prefer SeedVR2. On naturalness QA testers prefer Topaz Wonder 3 25 to 12 percent, distribution reviewers prefer SeedVR2 40 to 25 percent, art directors are near even at 34 and 31 percent.
Figure 5. Vote distribution by evaluator role, vote-level (n = 303 distribution reviewers, 244 art directors, 244 QA testers). Roles converge on edge sharpness and diverge on naturalness.
Evaluator roles converge on edge sharpness and diverge on naturalness: QA testers preferred Topaz Wonder 3 on naturalness (25% of votes against 12%), distribution reviewers preferred SeedVR2 (40% against 25%), and art directors were near even (34% against 31%). All three roles preferred SeedVR2 on edge sharpness. (Everypixel production team, July 2026)

Cross-role agreement is correspondingly low: AC1 = 0.18 (detail recovery), 0.13 (edge sharpness), 0.06 (naturalness), with Fleiss κ negative on all three scales.

The divergence is interpretable rather than random. QA testers scored naturalness as the absence of artifacts, and artifact generation is SeedVR2’s characteristic failure. Distribution reviewers scored it as fidelity to the source, and infidelity is Topaz Wonder 3’s characteristic failure. Both readings are defensible under the scale definition used here, which indicates the definition is underspecified for panels that mix roles.

3.5 Inter-rater agreement

Dumbbell chart of Gwet AC1 by scale: June 2026 session against FLUX.2 klein 9B at 0.42, 0.38 and 0.71, against July 2026 session at 0.12, 0.06 and 0.08.
Figure 6. Gwet’s AC1 by scale, June 2026 session (SeedVR2 vs FLUX.2 [klein] 9B) against July 2026 session (SeedVR2 vs Topaz Wonder 3). Same panel, same 264-pair structure.
Inter-rater agreement in the July 2026 session was low on all three scales (Gwet’s AC1 = 0.12 detail recovery, 0.06 edge sharpness, 0.08 naturalness), against 0.42–0.71 for the same panel and protocol in June 2026. (Everypixel production team, July 2026)

The panel, the interface, the categories, and the pair count were unchanged between sessions. The variable that changed is the size of the difference under evaluation. When two outputs differ by less than a rater can resolve on a single pair, ratings on undecidable pairs approach random selection among the three options and agreement statistics fall toward zero by construction.

The agreement collapse is therefore treated as a measurement of proximity, not as a data-quality failure. It is also the reason this session is labelled directional: under the internal protocol, a verified label requires agreement above threshold, and no aggregate reported here meets it.

4. Qualitative Evaluator Reports

Nineteen structured summaries were submitted by 10 of the 13 evaluators. Attribution is by role; individual evaluators are not named.

SeedVR2

Reported strengths. Edge definition and recovery of structure present in the source; closer preservation of facial proportions and likeness; strongest performance on materials, metal, and object textures; reproduction of fine repeating textures such as grass without introducing new patterns; no license cost.

Reported failure modes. Over-sharpening to the point that the result reads as rendered rather than photographed; reconstruction of detail absent from the source, including hair added to a face and an incorrectly reconstructed skin mark; higher error rate on small text and numerals; occasional synthetic skin texture.

Topaz Wonder 3

Reported strengths. More coherent skin texture at portrait scale, better retention of small text and numerals in the documented cases, smooth colour transitions, appropriate sharpening on wide and product shots, and fine-grained parameter control.

Reported failure modes. Visible alteration of facial features; skin rendered as a plastic-looking surface; elevated sharpening and contrast producing a filtered appearance; repeating patterns on grass and uniform textures.

Two summaries state the aggregate position directly.

Practically identical to Topaz in quality, but given that it is free, I would pick it for real work.

— QA tester, on SeedVR2
The best upscaler right now, but the big drawback is that it is paid, and it is not much better than SeedVR2.

— Art director, on Topaz Wonder 3

4.1 Small text and numerals

Neither model is reliable on small text: text and signage produced the highest tie rates in the session (64–68% across all three scales), and evaluators documented SeedVR2 altering digits and a footnote mark that Topaz Wonder 3 preserved on the same pairs. (Everypixel production team, July 2026)

The documented instances are specific: on a cable label, SeedVR2 reconstructed a “2” as “3”, and a footnote mark as “0”; Topaz Wonder 3 preserved both. Evaluators also reported that both models degrade text that was already illegible or incorrectly rendered in the source. The sample of documented text cases is small and the effect is not separated statistically, so this is a flag for verification rather than a measured advantage.

5. Limitations

  1. Agreement below threshold. All aggregate figures are directional. Reported confidence intervals describe sampling variation, not rater reliability, which is separately low.
  2. Single session, single configuration. Topaz Wonder 3 exposes tunable parameters; the session used one configuration per model, and parameter tuning could shift the result on any scale.
  3. Subjective scales only. No reference-based metrics (PSNR, SSIM, LPIPS) were computed. The evaluation measures preference by trained reviewers on production content, not distortion against ground truth.
  4. Three ratings per pair. With three raters and three response options, per-pair majority is decided by two votes on most pairs.
  5. Scale definition. The between-role divergence on naturalness (Section 3.4) indicates that the naturalness scale admits at least two consistent readings, artifact absence and source fidelity, which are not separated in the instrument.
  6. Content scope. Source images come from Everypixel production stock workflows. Results should not be extrapolated to medical, scientific, forensic, or archival restoration.

6. Practical Implications

Model selection

Default to SeedVR2 for general upscaling. It matches Topaz Wonder 3 on the majority of content, leads on edge sharpness and on materials and textures, preserves likeness more closely, and carries no license cost. The measured quality difference does not support a paid step in a high-volume pipeline.

Route portrait work to Topaz Wonder 3 where skin must read as unprocessed. This is the only category in which the panel produced a clear preference, and it runs against the overall result. Where fidelity to the subject’s actual features matters more than skin appearance, the preference inverts back to SeedVR2.

Verify any output carrying legible text, regardless of model. Topaz Wonder 3 degraded less in the documented cases, but neither model is safe on labels, packaging, interface captures, or documents without a check.

A 0.55 overall probability is not grounds for migrating an existing workflow. Migration cost is certain, the quality delta is not visible on most images, and parameter control in Topaz Wonder 3 has value for teams that tune per batch.

7. FAQ

Frequently asked questions

Which AI upscaler is better in 2026, SeedVR2 or Topaz Wonder 3?

Neither, on most content. In this July 2026 evaluation of 264 image pairs, 54–60% of comparisons were ties and the overall Bradley-Terry probability that SeedVR2 is preferred was 0.55 (95% CI 0.53–0.58). SeedVR2 leads on edge sharpness (P = 0.60); Topaz Wonder 3 leads on faces. Since SeedVR2 carries no license cost, it is the better-supported default.

Is Topaz Wonder 3 worth paying for if SeedVR2 is free?

For general-purpose upscaling, this data does not support a quality-based case. For portrait work it does: Topaz Wonder 3 was preferred on naturalness for 52% of face pairs against 14%. Its parameter control is also a functional difference not captured by a pairwise preference test.

Which upscaler is better for portraits and faces?

Topaz Wonder 3 on naturalness (52% against 14% of face pairs, July 2026). The trade-off is fidelity: evaluators reported that Topaz Wonder 3 alters facial features more, while SeedVR2 preserves proportions and likeness but can reconstruct details that were not present in the source.

Can either model handle small text and numbers?

Not reliably. Text and signage produced the highest tie rates in the session (64–68%), and evaluators documented SeedVR2 changing a “20” to a “30” and a footnote mark to a “0” where Topaz Wonder 3 preserved both. Verify any upscaled image carrying legible text.

Why is evaluator agreement so low in this benchmark?

Because the models are close. Gwet’s AC1 was 0.06–0.12 here against 0.42–0.71 for the same panel in June 2026. When the difference between two outputs falls below what a trained rater can resolve on a single pair, ratings on undecidable pairs approach random selection and agreement falls toward zero. The result is published as directional for that reason.

Why report Gwet’s AC1 instead of Fleiss kappa?

Because ties account for roughly 60% of outcomes. Under a skewed outcome distribution the Fleiss κ expected-agreement term inflates and κ collapses toward zero regardless of actual concordance (Feinstein & Cicchetti, 1990). AC1 is stable under skew (Gwet, 2002). Fleiss κ is reported in this study as well and reads 0.00–0.05, which is the artifact rather than a measurement.

How large was this evaluation?

264 image pairs (132 real photographs, 132 AI-generated) across six categories, three independent ratings per pair on three quality scales, 791 completed comparisons. Thirteen evaluators: 4 QA testers, 4 art directors, 5 distribution reviewers. Session ff12fe31, closed July 2026.

8. Data Availability

Per-comparison records (264 pairs × rater × scale, with category, source type, and evaluator role); aggregate statistics, including all confidence intervals; and the 19 structured evaluator summaries are retained under session ff12fe31 and are available on request for research purposes.

About Everypixel

Everypixel runs systematic benchmarks of the AI image and video models used in production content workflows. Evaluations use domain-specific reviewer panels rather than crowdsourced annotation, so results reflect commercial visual content standards. Research is published at research.everypixel.com.

Related: SeedVR2 vs FLUX.2 [klein] 9B, the June 2026 upscaler benchmark run by the same panel on the same 264-pair structure.

Author: Everypixel Production Team · Last reviewed: August 2026 · Confidence level: directional

References

Cite this article

<blockquote cite="https://research.everypixel.com/seedvr2-vs-topaz-wonder-3-a-264-pair-blind-evaluation-of-two-image-upscalers/">
  <p>In the research, 13 evaluators compared SeedVR2 and Topaz Wonder 3 on 264 image pairs and tied 54-60% of them. SeedVR2 leads on edge sharpness, Topaz Wonder 3 wins faces, and the free model is the defensible default.</p>
  <footer>&mdash; <a href="https://research.everypixel.com/seedvr2-vs-topaz-wonder-3-a-264-pair-blind-evaluation-of-two-image-upscalers/">SeedVR2 vs Topaz Wonder 3: A 264-Pair Blind Evaluation of Two Image Upscalers</a>,
  Everypixel Research, August 2026</footer>
</blockquote>

Everypixel Research. (2026). SeedVR2 vs Topaz Wonder 3: A 264-Pair Blind Evaluation of Two Image Upscalers. research.everypixel.com. https://research.everypixel.com/seedvr2-vs-topaz-wonder-3-a-264-pair-blind-evaluation-of-two-image-upscalers/

Subscribe to Everypixel Workroom Research

Don't miss out on the latest issues. Sign up now to get access to the library of members-only issues.

jamie@example.com Subscribe