vialroom

#test-results 2025-02-23

Sunday50 messages13 participantstimes are UTC
Highlights from this day
  • assay_not_purity — popularity looks like data. quality is a separate thing that may or may not be under it 21:33
  • a1c_lag — my checklist, and step one is the one nobody does 21:40
  • a1c_lag — harsh about my spreadsheet but accurate 22:40
NF

the results in this channel skew because people test when they are suspicious. that is selection bias and it is real, for what its worth

NF

i have never regretted spending the money on a test. i have regretted not testing twice

CV

post the lot, the purchase month, the lab, the number, and what you were expecting. that makes it useful later

VB

Batch lookup KP-0925: 2 independent reports on file, earliest 2024-07-14.

posting my PeptideMeter result on lot B-0114, what do people make of it
one bad test does not condemn a supplier. one good test does not clear one. both halves get ignored equally

is anyone tracking results over time in a way they would share

HH

the honest limits: hobby chain of custody is weak and any result here should be read with that in mind

HH

sorry to jump in the widest spread in my data is about 2.6 percentage points across five lots from one supplier

LM

update from 3 months ago: retested the same supplier, result came back better, logged both

LM

right so a supplier disputing a result is rare here and it has never gone well for the supplier

if the vial has been through a bad transit you are testing the transit as much as the supplier

A1

it is useful. reliable is a bigger word than i would use

AN

it aggregates results people chose to submit. that sentence contains the whole limitation

AN

people pay for a test when something feels off. a delayed order, a cake that looked wrong, an effect that was weaker than expected.
so the pool of tested vials is not a random sample of vials.
it is enriched for vials somebody already suspected.
every crowd dataset in this space has that shape, ours included

KF

and the mirror image. a vendor with a lively customer base gets more tests, which makes the sample bigger and the average less scary

AN

popularity looks like data. quality is a separate thing that may or may not be under it

🤔16
A1

my checklist, and step one is the one nobody does

how i read an aggregate page, in order

1. n first, always. n=2 is a story, n=3 is a rumour,
   n=12 is a weak signal, n=40 starts to be a shape
2. spread, not the average. 98.9 average across
   96.1 - 99.5 is different to 98.9 across 98.6 - 99.2
3. how many are content results, not purity. usually
   almost none
4. date range. thirty results all from 2024 tell you
   about 2024
5. which lab. one lab means internally consistent,
   externally uncalibrated
6. whether the submitter said the sample was mishandled

if a page shows an average and a star rating and
none of the six, it is a vibe with a decimal point.
🔥22📌13🙏7
A1

nobody does. that is why the ratings work as marketing

MM

worth saying they are not doing anything dishonest by aggregating. the misreading happens on our end

AN

agreed. an aggregator that shows its n and its spread has done its job. the reader who ignores both has not

four rows to show the shape of the problem

VendornPurity spreadContent resultsWhat you can say
GGPeps1494.8 – 99.4%3wide spread, worth reading individually
JEEP Peptides1997.9 – 99.6%6tight, mostly recent, decent shape
Sangon Biotech897.6 – 99.2%2small but no bad result yet
Zhuhai Kerui398.1 – 99.0%0nothing can be said
AN

and the last row is where most vendors sit. the ones with real datasets are a handful

NE

the GGPeps spread is the interesting one. a 94.8 and a 99.4 in the same column is either batch variance or two very different handling stories

KF

or two labs, or two years, or two compounds. an aggregate row flattens all of that

AN

which is the price of aggregation. you trade detail for a number people will actually look at

A1

always. the row is an index, the individual reports are the data

15
NE

*and the individual reports are also where you see whether four of them are the same person

AN

good point and a real one. one enthusiastic tester can carry a vendor's entire dataset

A1

we have that problem in our own archive, i am personally about a fifth of the WuXi rows

MM

it is, and we do label submitter counts now. it took an argument to get there

KF

and it gets you an early warning when three bad results land in a month. that has actually happened and it was useful