Put another way, the question to ask is not whether a vendor is good but what evidence exists, of what kind, about which lots, from whom. Reputation is a compression of that evidence and it compresses badly.
Lot-to-lot content variation of nine per cent between two nominally identical lots, both within a stated specification, is the single most common finding in independent testing and the least discussed. It is not fraud; it is the consequence of a fill process controlled to a tolerance rather than to a target. It is also the reason a per-lot content assay is worth more than a per-supplier reputation.
The difference between a batch certificate and a vial certificate is a difference in what is being claimed. A batch certificate says "we tested some vials from this lot". A vial certificate says "we tested this vial". Neither is worthless; only one of them is about the object in your hand, and the gap between them is a sampling assumption nobody has quantified.
The published aggregate datasets from Janoshik, Medutest and PeptideMeter are the closest thing to a systematic evidence base in this space, and the striking pattern across all three is that identity is almost always confirmed, purity is usually acceptable, and content is where the variance lives.
The limitation of the red-flag approach is that it is asymmetric: it identifies bad documentation reliably and good material only weakly.
If a supplier will not send you a lot-specific certificate before you order, you have learned something useful at zero cost.