Completely coincidentally, we're just about to launch a service that does exactly this (API access to uncensored open models)! We have a waitlist at the moment but will be live very soon!
There is very little information that is illegal by itself. At least in the Western World, and especially in the US. The question is how far you get into the territory of aiding and abetting a crime
But the reasonable defense is that the intended use cases are legal. The home page list a couple, and the 'writing fiction'/'helping authors' case alone covers almost everything. An author asking you how to best conduct a terrorist attack or how Meth is made are perfectly normal. Maybe even tame, compared to what some authors tend to research
More seriously though, I think we should be fine: we don't host any content, and what people do with the models is their own responsibility (legally speaking, in our jurisdiction, at least according to Claude -- we're talking to a real lawyer next week). Like any other provider, we offer no guarantees of sane, safe, or accurate results.
Thanks for bringing up a service like this, it's quite important. A few serious questions if you don't mind.
Confidentiality? Do you use any sort of logging and if not do you have a way to guarantee that your hosting providers are not snooping?
Price vs Vast or Runpod? If i have a very large or a very small workload do you have a competitive rate vs a gpu provider that offers private gpu access?
Subscription vs Api costs? Do you only offer api rate or will you offer discounted tokens for subscription? Subscription friendly towards open source harnesses such as omp?
Heretic ablation vs other methods? KL divergence scores? Do you post train the weights yourselves or do you offer weights trained by other organizations and is this information available on the service?
Cache hit/miss pricing policy? 90/10 or a different cache pricing policy, and how long do conversions stay in kv cache?
Quantized cache and model? Do you offer a choice if i want a quantized model for speed or a quantized cache? If not do you publish the information?
SGlang vs vllm or other inference engine? Do you publish your engine stack details?
Thank you kindly I find the competition in this space very lacking.
These are great questions, thanks! I'll answer them in turn, in a list because I like lists.
* Confidentiality: no logging, third party analytics, or anything like that. More details in our Privacy Poilicy [1]. Our hosting providers will have their own policies, but we're not running a super private service like Proton or similar. Might do some kind of secure tenancy in the future if there's demand.
* Price: I think Runpod vs per-token are very different beasts and for different purposes. I really can't make a direct comparison, as it'll be based on use case, but we're going for convenience over price, so all else being equal I'd expect us to be more expensive for most users anyway (edit: i meant "than other API providers"! We'd definitely need to be cheaper or at least competitive with spinning up your own cloud infra. We'd have parallelism and economies of scale on our side for this). We have a lot of experience with running and optimising open models though, so that's part of the value proposition too.
* Subscriptions: Only API for now. Maybe subscription later but honestly we prefer simplicity. My own experience with subscription plans is that they're usually sold at a huge loss at first, then the price creeps up as the service is enshittified. That feels like a bit of a scam to get users, and that's not really what we're about. We want to provide something specific, and aren't really concerned about scaling as fast as possible. Maybe we'll provide subscriptions if there's a real demand for it, but no plans at the moment to do so.
* Methodology: we use abliterated models, but I've been advised to hold off talking about that for now. Might make a blog post about this though (when we have a blog).
* Cache: yeah about 90/10 for pricing. We're still trying to find the sweet spot for tuning eviction. Running LRU with no guarantee/storage at the moment, could probably be less aggressive with retention, but that also has privacy surface area implications. Ongoing conversation.
* Quantisation: my brother in christ, everyone runs quantised. :) We're initially targetting FP8 on most models, but have had great results with MXFP4 though. If we can pack more concurrency onto nodes without losing quality, we'll reflect that in pricing. Or we'll offer as a separate model for cheaper and give users the choice. Edit: I see you were asking specifically about speed, which MXFP4 doesn't improve, but maybe if there's demand we'll run other qaunts for speed increase, especially on the larger models.
* Engine: vLLM gang all the way! For now at least, as it's what we have most experience with, and we find it the most flexible. We've been experimenting with SGLang though, and there's definitely some interesting optimisations we could do with it.
Hope this answers your questions, at least the ones I could! The irony of that hasn't escaped me!
Improper use is that of the user, not inherent to the tool.
Scolio: guns. Respondeo: guns are much more specialized (one-use) than knives. Proper use of sharp knives when what was shipped was a butter knife is understandable.
(The simile is not fully overlapping but should give the idea. The instrument must be flexible; if it is misused it is then a responsibility of the abuser.)
Generating blackmail is not illegal, using it to blackmail someone is. Generating libel is not illegal, publishing it publicly is not illegal either although you can be sued over it.
Generating worms and computer viruses is not illegal last I checked, but disseminating them is.
My point is that generating is not illegal but sending is. Your whole point hinges on legality, so the distinction between legal and illegal seems pretty key.
Its amazing how biased that source is and how much it buried the lede. They were arrested for participating in a riot where someone attempted to murder a police officer. That's not just someone criticizing ICE on social media. If it was, half the posters in any political thread on HN would already be arrested.
Would love to get a sub. Ive currently got three separate subscriptions to the models above and would be great to combine them.
Coincidentally, I'm Read teaming and checking security issues for a company with the same name as you.
This reminds me of an article [1] (that I think I may have even seen initially on HN) analysing what writing looks like when treated like TV (consciously or subconsciously). It's a fascinating read and it made me more mindful of interiority in my own writing.
That AI consumes it at an abnormally high rate, presumably.
The claim never made sense to me either, I can only assume those that regurgitated such claims never worked with HPC or even general datacenters before.
Was recently talking to a (non-technical) friend about this, she was surprised after talking about the "insane water use for AI datacenters" when I responded that open-loop cooling is pretty rare for a datacenter and I've never actually seen it used before, versus closed-loop (or just regular air-based cooling) which has no real noticable water consumption.
Problem with data centers is that companies want to build them near densely populated areas that already have problems with water supply and high utility bills.
1. They don't HAVE to use water. Air cooling, closed loop cooling, waste-water cooling, and so on, are options. Easy to regulate. Evaporative cooling is more energy efficient though, but a complete non-issue in places with abundant water and a non-option elsewhere.
2. Datacenters have been shown to reduce utility prices. They provide suppliers with previsible long term demand which allows for cost-effective network and production planning.
That AI data centers are drinking up local ground water for cooling. It isn't (or wasn't) nonsense, though. It was/is a real thing, though it seems to be on the out in favor of closed loop cooling after the massive and still on-going public outcry.
Because, obviously, you should be spending all of your waking time thinking about LLMs, agents, and how you can integrate them into every part of your life. If you have been living properly in the age of impending-AGI, you would have already been desperately seeking more opportunities to interact with these systems. That desperation would have led you to independently discover agents and all the ways you could couple yourself to them even when away from your computer. Are you a parent stuck at home experiencing life with your kids instead of sitting at your desk? Why not escape such a hellscape by whipping out your phone and building a SaaS from your phone while your offspring annoys you with requests for attention and meaningless affection?
---
Really, this whole environment of 'coding from my phone with dozens of agents while I'm doing the laundry' feels like satire of the sorts of things we used to laugh at on Linkedin.
The closest I've ever felt to this as a native English speaker is reading words in music scores in English. I'm a classically trained cellist, and grew up learning notation with Italian and French words for directions and expression. I've never learned either of those languages, save the words used in music notation. Seeing a score with those words in English just feels... wrong. Not in any big way, but as you said: uncanny. Definitely get the "bad psuedocode" vibe, because to me it English in music notation feels similar -- like the person who wrote it didn't know what they were doing, even though the notation makes perfect sense and the music is good. It removes some of the flair of the art of the notation itself for me.
This is a very good analogy, sheet music with all the Italian replaced by English would be very funny. "Loud!" "Very loud!" "Super-duper quiet!" "This is the end of the song!" "play this part reallll smooootthh" etc.
(Actually I believe the late P.D.Q. Bach [1] did this a lot, and it was in fact quite funny)
This reminds me of an article I read that was posted on HN only a few days ago: Uncertain<T>[1]. I think that a causality graph like this necessarily needs a concept of uncertainty to preserve nuance. I don't know whether this would be practical in terms of compute, but I'd think combining traditional NLP techniques with LLM analysis may make it so?
I get some vibes of fuzzy logic from this project.
Currently a lot of people research goes in the direction that there is "data uncertainty" and "measurement uncertainty", or "aleatoric/epistemic" uncertainty.
I foumd this tutorial (but for computer vision ) to be very intuitive and gives a good understanding how to use those concepts in other fields:
https://arxiv.org/abs/1703.04977
Right. The first example on the site shows disease as a cause, and death as an effect. This is wrong on several levels: There is no such thing as healthy or sick. You’re always fighting off something, it just becomes obvious sometimes. Also, a disease doesn’t necessarily lead to death, obviously.
Since you're always going to die, the problem is solved - the implication is true by the right side always being true, and the left side doesn't matter.
I don't think the part about front and back channels is quite correct. GET and POST requests are both encrypted in HTTPS -- including the URL (but not the domain, as DNS resolution happens separately). Front and back channel are more to do with trust boundaries, and what information is public vs private from the client's perspective.
You can't. It is just a matter of reducing the risk surface. With a GET someone may add parameters, with a POST they would send the data in the post (which is often the main point of a POST).
Since all typical web servers/processors only loh the call and not the body there is a lesser probability of a leak.
I am writing this as someone who manages cybersecurity and is offering faced with not enough information in investigations because of that. This is also the reason that I used "typical" and "usually" above - it is pretty weird what people send and how they process what they receive.
If your experience of tofu is only the above, I completely understand your distaste for it. But I think you owe it to yourself to try better tofu, and not as a meat alternative. Tofu on its own doesn't have much flavour, but that's the point, you need to marinate it. Google can give you some tips (squeeze the water out, then soak it in something delicious -- hell, soak it in meat juices!), but I highly recommend trying some good, low moisture smoked firm tofu. It's so good, I often just snack on it, slicing it like a sausage. But it's also great in things like burritos to add a smokey kick. Try it!
https://violentdelights.ai