Or just have the last data point include everything from the full 12 month period before (as the title "year on year" would suggest) and maybe even put it in the correct place on the x-axis (e.g. for today, Feb 12, 2025, about 11.8% of the full year gap width from the (Dec 31) 2024 point).
IME Netflix is a close 2nd best after Apple, which I don't think I can distinguish from a 4K BluRay. I've found that the quality depends on the platform a little -- for Netflix the native LG app seems to look best on my LG TV, while Apple looks best on the Apple TV app (perhaps unsurprisingly).
Amazon Prime 4K HDR on the other hand looks like garbage on every platform I've used -- the compression is unbearable in any dark scene.
I would put Disney+ after apple. Both AppleTV+ and Disney+ consistently looks great to me. Netflix is strange as it generally looks good but whatever compression they use does something funny to the picture which makes it look fuzzy and sharp at the same time to me.
Netflix: 15-18 Mbps
Disney+: 25-30 Mbps
Amazon Prime Video: 15-18 Mbps
Apple TV+: 25-40 Mbps
HBO Max: 15-20 Mbps
This is from an LLM but it tallies with what I remember reading. Apple TV is by far the best, followed by Disney+.
Netflix unfortunately seem to use any improvement in compression encoding efficiency to reduce bitrates, rather than improve PQ at the same bitrate. It's definitely got worse over time. I also remember reading that for content they deem more compressible they use a lower bitrate.
I can sort of get that on the lower plans, but its frustrating they won't improve PQ (or at least keep it the same) for the (expensive) 4K plan.
I like JAX but I'm not sure how an ML framework debate like "JAX vs PyTorch" is relevant to DeepSeek/PTX. The JAX API is at a similar level of abstraction to PyTorch [0]. Both are Python libraries and sit a few layers of abstraction above PTX/CUDA and their TPU equivalents.
[0] Although PyTorch arguably encompasses 2 levels, with both a pure functional library like the JAX API, as well as a "neural network" framework on top of it. Whereas JAX doesn't have the latter and leaves that to separate libraries like Flax.
Erm, why not? A 0.56 result with n=1000 ratings is statistically significantly better than 0.5 with a p-value of 0.00001864, well beyond any standard statistical significance threshold I've ever heard of. I don't know how many ratings they collected but 1000 doesn't seem crazy at all. Assuming of course that raters are blind to which model is which and the order of the 2 responses is randomized with every rating -- or, is that what you meant by "poorly designed"? If so, where do they indicate they failed to randomize/blind the raters?
> If so, where do they indicate they failed to randomize/blind the raters?
Win rate if user is under time constraint
This is hard to read tbh. Is it STEM? Non-STEM? If it is STEM then this shows there is a bias. If it is Non-STEM then this shows a bias. If it is a mix, well we can't know anything without understanding the split.
Note that Non-STEM is still within error. STEM is less than 2 sigma variance, so our confidence still shouldn't be that high.
Because you're not testing "will a user click the left or right button" (for which asking a thousand users to click a button would be a pretty good estimation), you're testing "which response is preferred".
If 10% of people just click based on how fast the response was because they don't want to read both outputs, your p-value for the latter hypothesis will be atrocious, no matter how large the sample is.
Yes, I am assuming they evaluated the models in good faith, understand how to design a basic user study, and therefore when they ran a study intended to compare the response quality between two different models, they showed the raters both fully-formed responses at the same time, regardless of the actual latency of each model.
I did read that comment. I don't think that person is saying they were part of the study that OpenAI used to evaluate the models. They would probably know if they had gotten paid to evaluate LLM responses.
But I'm glad you pointed that out, I now suspect that is responsible for a large part of the disagreement between "huh? a statistically significant blind evaluation is a statistically significant blind evaluation" vs "oh, this was obviously a terrible study" repliers is due to different interpretations of that post. Thanks. I genuinely didn't consider the alternative interpretation before.
Sure, it could be, you can define "preference" as basically anything, but it just loses its meaning if you do that. I think most people would think "56% prefer this product" means "when well-informed, 56% of users would rather have this product than the other".
Isn't the "ban the Apple-Google iOS search deal" just one of several proposed remedies, with the most significant one being a Google breakup? Certainly seems like that one would affect Google more than Apple. Or am I confused and the Google breakup thing is a proposed remedy in a separate case?
I assumed if I kept reading there would be a line explaining why they can't simply raise prices until the demand becomes manageable with current staffing, such as "We sold all these suits with guaranteed adjustments for £[some heavily discounted number] for life", but I didn't find any such explanation. Shrug
I think the population of people buying bespoke suiting is small enough that you would not want to alienate your existing customers. I agree that they should raise the prices, but I've got to think there's an aspect of a relationship there. It was hinted at, a little bit, in the article. It's not just a financial transaction, I mean.
Precisely. They're talking about a customer who has spent £700,000 ($870,000) on suits. That's a long-term relationship built on trust. Hiking your prices to manage demand might be a short-term financial bonanza, but it's disastrous in terms of reputation.
And the article suggests that's it's not even the population of everyone with a bespoke suit so much as the minority of whales who own a lot of them. There is going to be a fair number of very demanding and impatient rich guys in that group.
Yes, this is correct -- TikTok's own "shutdown" was never required by law. I'm not in the US so I can't check for myself, but from googling it still seems to be removed from both the Apple and Google stores.
If Apple/Google don't change their minds, TikTok won't be able to get any new US users, and won't be able to distribute updates to current US users. To continue using it in the current state, US users will have to keep the same phone and TikTok will have to continue supporting whatever last version(s) they're on indefinitely. (Modulo the few that might jump through VPN and app store locale setting hoops.)
And I don't see how Apple/Google could change their minds: the ban bill comes with a 5 year statute of limitations, so regardless of how convincing the Trump administration is in their promise not to enforce the law, the next administration inaugurated in January 2029 would still be able to impose the penalties on Apple/Google for 4 years of non-compliance. Those penalties would be cripplingly massive even for the world's largest companies (I'm reading an estimate of $850B [0]).
As far as I can tell, the only events that could end this are (1) TikTok finding and agreeing to sell to a US buyer or (2) Congress overturning the ban.
It's odd that people are talking as though the saga is over now...
Yes, I'm so completely fed up with recurring subscriptions for things with negligible or no recurring costs for the seller. This one is of course particularly obnoxious given the hardware itself is expensive and the recurring cost is 0. But for example I would've gotten an Oura ring by now if they would just charge twice the price for the ring itself and not require a subscription, even though the subscription fees over the lifetime of the hardware would probably add up to a significantly smaller amount. To me it's just incredibly off-putting -- it reeks of greed and feels like a blatant attempt to fool customers by obscuring the actual cost. I guess it must be working for them, but for me, the cost of anything with a recurring fee gets mentally rounded up to "approximately $infinity".