If you hook an Apple TV (or similar) up to any 4k TV, and turn CEC on (to let the Apple TV control it), and never set it up otherwise, it becomes a dumb monitor.
I got a degree in mechanical engineering. There's these two technicians under me that for years I thought we just hired off the street and maybe they graduated high school. I have to show them how to use terminal. I have to show them how to fix a Python script.
Imagine my horror when I found out recently both of them have a computer science bachelor's degree from a UC Santa Cruz.
They were asking why the two deepest holes, despite being nowhere near each other, dug decades apart, are 99.3% of 12km and 99.5% of 12km respectively. Was BP symbolically honoring the russian scientists? Does the earth have an extremely uniform material property that happens to be at a very round number of km? Just a complete coincidence all around?
(I asked AI, and it says coincidence, since BP stopped drilling once they hit oil, and the russians stopped drilling once they hit some melty rock.)
Also in both cases economic reasons. BP drilled to reach oil which makes economic sense, but AIUI the Russians wanted to keep drilling but eventually central government wouldn't give them any more money.
But yes, largely a coincidence. I think humans would see the same "pattern" if it were slightly more than 12km. We like patterns, we're the superstitious pigeon experiment but at a ludicrous scale. I would like to think the patterns I've seen point at some underlying more important truth - but the pigeon thought so too.
Step back and think about it another way - ask which scenario is more likely:
Some random person discovered a 60% across the board gain in all LLMs, using an extremely simple trick that none of the labs noticed in all these years. That trick being to rasterize 8bit characters into 8x8 pixels in a big image. 60% in a market worth trillions of dollars.
or
Anthropic's marketing team arbitrarily prices tokens to drive growth, according to vibes and feelings, and didn't think they needed to price images on par with text in their rush to burn cash & drive growth. Some folks take advantage of the trick during the first few days of the model's availability before Anthopic corrects their pricing, to align more proportionally with actual compute costs.
Nah, optical compression is a thing. You see it in a lot of different areas in ML. In this case, the "trick" has been known for a while, and belongs to a whole world of compression research. But I think where you're maybe getting mixed up is in where that 60% gain is coming from.
It's not a 60% percent reduction in cost for 100% of the same output. If you have a model and input text A, and you fix the seed etc. and run Text A through the model as text tokens and as compressed image tokens, you will not get identical outputs. You're specifically reducing the number of tensors needed to represent your input, which saves you on raw compute, but also by definition gives you less room to represent the information in your input. It's lossy, in other words.
Put another way, if you're using a model like Fable because you need the absolute frontier of capability and cheaper models cannot solve your tasks, then there is a very real chance that a compression strategy like this drops Fable's accuracy such that it's no longer suitable for your task. Which defeats the point of you paying for the most expensive model in the first place.
So, it's cool research. Might be useful for some people. Probably isn't something that has incredible utility in real use cases.
To me compression implies smaller size? However new line chars seems to be removed in the pic so I guess it could be expressed in fewer bytes than the original text with further compression ...
The size is indeed smaller, because text tokens and image tokens are embedded as vectors of the same size, but text tokens typically only cover a few characters, while image tokens typically cover many pixels, so many that you can fit more characters in there. So the same text takes up fewer tokens as an image, and hence requires less time and memory to process.
You could also imagine models where text tokens cover many characters and image tokens just a few pixels, which would invert the relationship, but this is typically suboptimal for the applications people have in mind when they train a model.
Lots of researchers have done just this! There's a really rich history of research + lots of contemporary work on different encoding/representation strategies. This might be interesting to you: https://sbert.net/
What makes the DeepSeek-OCR and related results exciting to some researchers is less about the fact that you could devise a tokenization scheme that has fewer tokens, and more about how well it works.
> Some random person discovered a 60% across the board gain in all LLMs, using an extremely simple trick that none of the labs noticed in all these years of multi-trillion dollar growth
DeepSeek published a pretty well circulated paper on exactly this many months ago. It just hasn’t been attempted and shared publicly, asa retrofit, AFAIK.
Also, it’s no free lunch, the readme indicates that this “use images” hack is lossy and reduces success rates alongside the reduced cost. Most labs would focus on success increases regardless of price.
If the trick were genuinely useful, and was well circulated months ago, the resource-starved inference providers would have squeezed this trick dry already, instead of wasting 60% of their tokens, waiting for users to implement it themselves in 5 minutes of effort.
That's like saying quantization isn't real because the frontier labs aren't using it in their production inference.
This is a lossy process, it produces worse results. It might be worth it for some situations, but applying it to everything would just be making your SOTA model worse
The "trick" is well documented in their Deepseek-OCR paper, that builds on plenty of other work. It's just not simple to just switch a commonly used LLM architecture to a new one, but I don't doubt most frontier labs are already experimenting with it. This by itself a very active field of research.
Isn't this just quantization with extra steps? Can converting the text to an image really be a better way to lossily compress it? (Not that I have any idea what I'm talking about on this topic.)
I also have no idea what I'm talking about, but to me this seems closer to the "caveman mode" that some people use to compress info into fewer tokens. Going through the image tokenizer allows you to leave the source text untouched while still gaining (some of?) the benefits
No, quantization is applied to model weights or the KV cache (the model activations of all past tokens), and is just storing everything with lower precision (carefully, so that it doesn't hurt performance much).
Sending an image of text instead of text reduces the number of input tokens, but they're still being processed by the model at the same precision. This probably also hurts performance in some way – the question is by how much.
I think you missed the part where this is a lossy technique that reduces performance.
The image trick reduces context because it’s lossy. The README says you can’t use it for anything needing exact recall. It produces a gist of the input.
You could achieve something similar by using a small, cheap model to pre-summarize information for the expensive LLM. This is what many people do already and it’s a much better way to do it for most situations.
This has been known since VLMs were a thing, that more information can be encoded visually and token efficiency is increased. But it came with performance issues (more hallucinations, etc).
Also I don't think you realize how much dumb stuff is still left on the table. That the market is worth trillions is quite irrelevant here given the dynamism of the field.
Alternative 1 isn’t all that unlikely given Opus 4.8 couldn’t do this. So it’s a recently possible hack. Not something LLM corps have been blindsided by for years. I also strongly recommend RTFA in this case, namely ”The honest part, read before relying on it”
As long as we're nitpicking every sense of the word "own", the strongest legal sense means you're the copyright holder, and every sense downstream of that is some lesser license. Buying a disc is a license to view the intellectual property, subject to various restrictions like only showing it within your personal home.
If the disc is an abstract license, surely the seller will replace the disc if it's scratched. I already bought the license, so what is the real purpose of the physical token?
Somehow the concept of ownership has been twisted to so that obligations only flow in one direction. Rules for thee, not for me.
The point OP is making is that it's not the concept of ownership that has been twisted, there just never was ownership of media beyond owning the actual copyright. Everything else is licensing.
> various restrictions like only showing it within your personal home
Are you implying that lending the disc to a friend so they can watch in their own home is forbidden? Or taking the disc to the friend's place to watch together?
No, those aren't the restrictions. But there are restrictions. First-sale doctrine allows lending. But you are not allowed to play the movie in, say, a restaurant, theater, or other public place.
I understand your snark, but my I wasn't attempting to be pedantic for the sake of it.
Your original message said "buying a disc is a license to view the intellectual property, subject to various restrictions like only showing it within your personal home". This biases the interpretation towards more restrictions than rights, but I don't think that's the case at all. The restrictions are essentially - don't make money out of playing it, don't play it in public. Pretty much anything else goes.
Call all your friends and family to watch with you whenever you want. Lend to your friends. Take it to your friend's house to watch together. Sell it to someone else. Watch it as many times as you want, anywhere and with anyone you want - as long as those restrictions above are respected, which isn't hard to do.
Do I think this is ideal? No, maybe not - but that's the world we live in, and all things considered, that's still substantially more rights than most "digital ownerships" give you.