A difference worth investigating, not a verdict
Illustration: at the same 45-day observation point, size M has 20 received returns from 100 delivered units (20%); size L has 18 from 60 (30%). The difference is 10 percentage points, not 10% relative. L’s rate is 50% higher relative to M, but these counts alone do not establish a fit defect or statistical certainty. Check comparable styles, reasons and pending returns before designing a fit-information test.
Define the unit and denominator
For this review, use physical units returned from a delivered-item cohort divided by the original delivered units in that cohort. Name the market, style, colour, size, delivery period and observation cutoff. Do not divide returned units by orders: one order can contain several items. Exclude cancellations before delivery from this denominator and report them separately.
Define when a return counts: requested, received or accepted after inspection. Use one status consistently and keep pending cases visible. A refund can occur without a returned unit, and a returned unit can be exchanged rather than refunded. These states answer different operational questions.
Join the original item before comparing sizes
Build the join around order line and variant identifiers, not a product title that can change. Keep size and colour as sold at the time, delivered quantity, return quantity, relevant dates and reason status. Aggregate only the fields needed for the review; a size-level brief does not need shopper names, addresses or email addresses.
For an exchange, retain the original outgoing unit, its return and the replacement as linked events. State whether the replacement enters a separate cohort or is handled in an exchange analysis. Do not silently erase the original return because another item was shipped. Conversely, do not count both a return request and its receipt as two returned units.
Give cohorts equal time to mature
An item delivered yesterday has had less opportunity to return than one delivered six weeks ago. Choose a window suited to your actual return policy and operational delay, then compare cohorts at the same age. The window is an analytical choice, not a universal benchmark. Show pending and late returns alongside the settled count.
Calendar sales and refund reports can place the sale and adjustment in different periods. Do not read a high refund week as evidence that this week’s new buyers return more. Match back to the originating items for the cohort question, and preserve the calendar view for the separate cash or operational question.
Investigate reasons without inventing certainty
Show counts beside percentages and retain an “unknown” reason category. Review whether reason codes are customer-selected, staff-entered or inferred; they do not have equal evidential weight. Compare the product measurements, size guide and relevant service feedback. A small cohort, a promotion or customers ordering multiple sizes can affect the pattern.
Use this at work
- Define returned units and original delivered units.
- Join original order lines, variants and return events.
- Compare the same cohort age and report pending cases.
- Review counts and reasons before proposing a fit change.
Reference material
Platform guidance checked 3 October 2026. Examples and working checklists are Faccelerate editorial illustrations.