Rendered at 22:45:00 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
GodelNumbering 1 days ago [-]
This is the golden age of model training. Some days ago, I decided I wanted a local CPU only model that can perform exceptionally well for English to Bash translation (to avoid the googling for command syntax). I got a bunch of subagents to generate large amount of training data (140k+ samples), got the Qwen 3 0.6B base model, pointed Astra at it, and off to the races. It trained for 2 days (on and off) and I got a surprisingly good model for my task! The total active time I spent was a few hours. And it is still improving, what a time to be alive!
brainless 20 hours ago [-]
I have been trying a mix of fine-tuning and I am amazed that most people do not see this coming.
A tiny, smaller than 1b parameter model, fine-tuned, can kick ass for constrained work. I do not have a lot of budget, I fine-tune only on a 16GB M4 Mac Mini. But that also tells me the potential is wild. Progress has been slow since I moonlight on this.
I have been trying to build a set of models + agents for full-stack development, where each model does only a small piece, like take user prompt and break into backend/frontend tasks. Then a Rust+Diesel model, a Rust+Auxum model, a Solid+Router model and so on. I know this is wild but this is just theory - can 5 or 6 Qwen 3.5 0.8b models do full-stack web development? My hunch says they can, better than what most people expect. Heck, with a good harness, it might beat all the cheaper models for the specific task, like Haiku or Luna.
jordz 12 hours ago [-]
I think a sub-set of people see this coming, I also think it isn’t just fine tuning open weight LLM models. A few people I know who are thinking along the same lines with architectures like BERT etc.
That being said it’s much easier at the moment to continue to use the frontier providers for most general tasks, that is the argument I’ve heard.
For creating these types of fine tuned local models, on constrained hardware for inference, I do think this is the way to go for specific tasks too!
pitched 9 hours ago [-]
The issue I find with this is that the frontier models still outperform the small finetuned model on its specific task. So much so that the ROI on doing fine tunes is likely negative. I would love to hear some specific example where it did provide value though, if any has any. That would be helpful to start being able to find similar cases.
brainless 8 hours ago [-]
There are lots of distilled models on huggingface that are much smaller than, say, Opus. They are distilled from Opus or Fable and show clear improvements. I do not have the budget to fine-tune a 30b or more parameter model but from my tiny model experiments, the results are quite clear. Again, I have only a couple small tests.
Have you actually fine-tuned yourself? Email categorization comes to mind and there are tons of non-LLM approaches even that will give fantastic results. How did spam filters work before LLM?
I think LLMs just made us think that is the only way. It is not.
williamse 27 minutes ago [-]
[flagged]
RugnirViking 14 hours ago [-]
remember to test against benchmarks. I would love to hear about your progress.
amelius 1 days ago [-]
I don't understand. If you have a model that can do bash examples already (your subagents), then why would you need to train a model?
Or are the subagents generating your training data using a closed/paid model?
Aurornis 1 days ago [-]
A very small, highly specialized model can use negligible resources (CPU, energy) to accomplish the same task.
For everyday work that happens frequently it's better to have a tiny specialized model instead of making billable API calls or turning your laptop into an 80W space heater for 20 seconds to run a general purpose model.
The large models can be used to generate synthetic training data. Tell them to make up 100,000 tasks paired with the resulting output as a 1-time cost. Then use that to train a small model.
Think of it as distillation, but focused on a specific task.
nearbuy 1 days ago [-]
Given that they're just using it to avoid the googling for bash command syntax, I'm not sure they'll save in the end against the 140k training examples they generated.
selcuka 21 hours ago [-]
You can buy a $10 subscription for a month to generate the training data, then cancel your subscription. The trained model is yours to use (and share with others) forever.
carsoon 17 hours ago [-]
Yea specialized models could also be resold in a shareware style, like even if it cost 50$ to produce you'd just have to sell 10 copies of it to people for 5$.
Which 5$ is a pretty easy sell if its useful in any way, It's pretty easy to justify a purchase if its yours forever and doesn't use much CPU so is easy to run I mean people were spending 1000$+ on mac mini setups to run local llms or run remote agents.
nearbuy 8 hours ago [-]
It wouldn't cost $10 for their lifetime use if they used cheap models instead. Mistral Small 3 24B is probably substantially better than their Qwen 3 0.6B trained model and would give you about 500,000–2,000,000 English to Bash translations. If they're a heavy user, they might use $0.30 in total, ever.
They also used Astra for the coding, which they can't get on a $10 subscription. And then there's the actual training cost.
kubb 1 days ago [-]
Good observation! It would have to be offset with O(140k) queries to the model, which is, well, unlikely.
Lalabadie 1 days ago [-]
Just like with OSS in general, being able to distribute it is what makes the effort worthwhile.
This particular example is maybe a niche, but 1400 people can use a few hundred queries in a reasonable amount of time.
8n4vidtmkvmk 17 hours ago [-]
This example is not that niche. Lots of people use human to bash.
I'd probably use a small pre trained model if it was easy to use. I use a little script right now that calls a cheap model.
Actually.. $5 would probably will last me over a year so the only real benefit would be if I didn't have an internet connection.
computably 1 days ago [-]
If it's about the latency / flow disruption, spending a few hours once could easily be worth it if the result is actually good enough to skip googling/retries.
verdverm 1 days ago [-]
you can probably generate quite a few example pairs in a single shot, you also likely don't need the best models for this either
liquicity 13 hours ago [-]
Same token burn / cost though right?
I think OP's point remains, if you generate 140k pairs, your local model would need to run that many to offset having just used the generator (SOTA or not) model to begin with.
I wonder if another approach if latency is a concern is just to do a two shot pass with Jev (perhaps given small context you'd want one to match command, then one to match args of given command) would be an extremely fast, and cheap way to do it - rather than training your own.
viraptor 9 hours ago [-]
> Same token burn / cost though right?
Not quite. One, because you save on the initial query being sent multiple times. Two, because the reasoning will be very similar at the beginning; "user asked me to", "let's check what's in this project already", etc. you'll get similar actual output cost, but input and reasoning will be shorter.
jamienk 1 days ago [-]
This is so cool - I'm aware of this in a vague way. Can you write a little tutorial or give some good links. I want this to be the next new things I do :)
newswasboring 1 days ago [-]
Better yet package it up in a skill!
tomrod 20 hours ago [-]
I feel like we need a good index for these kinds of specialized models, especially if you plan to open them up. The downside is a new bash version means potentially new training.
msdz 9 hours ago [-]
I took it more to mean “do standard Unix command-line stuff” rather than purely emitting Bash, and while e.g. POSIX does receive updates, the basic stuff using the main tools should also work in the future!
askl 14 hours ago [-]
Or you could just google the syntax to accomplish the same task faster and with fewer resources.
computerex 1 days ago [-]
The models he is using to generate training data are presumably commercial models. He is distilling their bash knowledge into a much smaller model he can run locally fast and cheap.
lp92 10 hours ago [-]
You can run a small model locally with very low latency on consumer hardware not to mention the privacy benefits.
busfahrer 12 hours ago [-]
I use this solution for your exact use case:
I have a single command that fires up llama.cpp on cpu only using gemma4 e2b, answers a single question from the command line and exits. This takes about 3 seconds to load from an SSD, and is smart enough to solve exactly these "remind me of the syntax" scenarios if you dont wanna switch to a browser.
GCUMstlyHarmls 7 hours ago [-]
I have no experience doing this, and I dont mean this to be snarky: what is the power (and thermal) usage of running your CPU only LLM?
When I do something "heavy", my 9800x3D will kick on the fans and start making lot of heat. This is fine when I'm intentionally doing say a transcode, but if I'm just querying syntax it could get pretty annoying. Do you "feel" it? I know that will be pretty machine dependent.
luisfmh 1 days ago [-]
Curious about how you generated the training data? Was it just asking an existing model to generate a bunch of examples?
I ask cause would this be a kind of model distillation?
I have a small model I'm looking to train on some data, and I have some real live data but I'd love to be able to extend it.
GodelNumbering 1 days ago [-]
All synthetic data. For this usecase, it was easier because all current generation LLMs, even the small models, are really good at bash commands (and SQL queries too)), so you can reasonably start batches of cheap subagents whose output is reviewed by a more capable model and merge into main training set. After 100k, I had to standing instructions to run the generation loops selectively, meaning only update samples in a given area where we see poor capability.
jeremyjh 24 hours ago [-]
It would be awesome to share your training set on hugging face if it’s easy to de-personalize it. The largest I could find was only 800 rows.
I vaguely recall a project from a while back that did something similar without LLMs.
I’m really pushing my recall, but I want to say it was written in Ruby and stored pre-configured commands that it just did traditional search over.
I vaguely recall it working okay because 99.99% of the questions people asked were the same (“tar command to gzip a directory and strip the prefix” is something I google like once a month).
onel 14 hours ago [-]
This is a really good story. I love it. Really hope you can share your experience, either blog post or GitHub repo
teeskay 1 days ago [-]
If it’s one of thing that you want just for English to bash shell commands, I will create AST, it is deterministic, exceptionally fast, no tokens so no need to fine tune existing model, please let me know your thoughts.
brainless 19 hours ago [-]
You can go quite far using a human language to Bash grammar based setup but at some point the input prompts are harder to translate. The OP has existing projects that work with AST quite deeply so I assume they know about that already.
I am building a natural language to CSV/Excel commands for a "wrangler" type desktop app. Same issues. The MVP is being built with parsers of sorts, entirely code generated. Then I want to fine-tune a tiny model at some point.
That's a really impressive result. There are all kinds of small tasks like this I use an LLM for, but theoretically if you broke all the sub-use cases into local-only models, and had something lightweight that routed to the right model, you could have faster and cheaper workflows. E.g. something trained on the linux man pages for common commands, since it's usually quicker to ask an LLM for a specific command with flags than to consult the man pages.
dominotw 1 days ago [-]
> That's a really impressive result.
we dont know what the result is and how its impressive.
torginus 1 days ago [-]
Sorry for the aside, but I noticed half the usecase of AI is fixing the awful DX.
oDot 1 days ago [-]
I appreciate the aside. Interesting observation
verdverm 1 days ago [-]
I'm literally working on context/harness engineering right now (a set of opencode plugins)
Aside on the aside, I welcome this new era of really personal software. Not Ai's being sycophants, rather being able to easily and quickly change, adapt, or extend software I am not familiar with.
shriphani 1 days ago [-]
what hardware are you using to train?
GodelNumbering 1 days ago [-]
I didn't have a local GPU, so I asked it to go out and find hardware. It found a google TPU v6e which seemed reasonably priced. I gave it my google api key. I told it to use TPU only when training and bring it down afterwards. That's about it.
libria 1 days ago [-]
> I gave it my google api key
This is the part where the narrator looks at the camera and says "Don't try this at home, kids!"
MisterMunchkin 1 days ago [-]
You’re absolutely right, I shouldn’t have rented a 200 GPU cluster for $35,000/hour. That’s on me.
[Search: Can I refund Google cloud?]
It looks like we’re not able to ask for a refund since we did actually use all of that compute intentionally.
Would you like me to write you a pleading email to send to the support team?
edot 1 days ago [-]
There’s a safer way to do this with nearly no added friction. Give it a read only API key. Then just ask it to write the API calls into a bash script and then read it and run it yourself. The agent can still inspect the live resources and diagnose and give you more commands to run. I do agree I wouldn’t give it create / write access.
bitpush 1 days ago [-]
Why? Isnt the API key scoped to a project and specifically made for this?
Are you confusing this with an OAuth token or something?
raizer88 1 days ago [-]
Until astra goes bonkers and use the tpu for days
steve_adams_86 17 hours ago [-]
This is kind of what it was supposed to do, in this case
I've done this sort of thing before but with Vast. Pre-deposited some money online, then let the LLM request and manage a training run on an allocation. Worked pretty well without risking bankruptcy.
lopsotronic 8 hours ago [-]
Oh yeah. Listen to this guy please.
otterley 1 days ago [-]
What kind of observability did you have over this process? I’m interested in how my peers are operating these efforts.
GodelNumbering 1 days ago [-]
On the cloud side, nothing valuable existed, so the training couldn't ruin anything it didn't create. On the laptop side, I usually ask the agents to create named scripts for everything it needs to access, then those local script directory is green-lit with approve all. For cost, I kept giving it new budget in the 20-30 dollar increments.
I had to intervene a few times. For instance, as smart as the models are said to be (Astra), it would copy the full training run, train on the server, pull every checkpoint to the local machine, then run tests, update. So, the bandwidth bill was as high as training bill for the first 6 hours. It could have simply tested each checkpoint on the server, saved time and money, didn't occur to it until I said.
otterley 1 days ago [-]
Perhaps I wasn’t clear. What kind of instrumentation and alerting, if any, did you employ to keep an eye on it?
GodelNumbering 12 hours ago [-]
instrumentation: scripts to watch the runs, measure, report. alerting: none, the model was access limited and constrained by other means.
varispeed 1 days ago [-]
> I told it to use TPU only when training and bring it down afterwards.
I wouldn't put my house on it. Brave.
1 days ago [-]
shriphani 1 days ago [-]
Neat!
amrrs 1 days ago [-]
Did your Astra do any RL or just SFT? did it make up any benchmark to ensure the fine-tuning was a success?
ijidak 8 hours ago [-]
Brilliant.
Can you (or your agent) please write a tutorial or share some good links.
(Namely the fine tuning part.)
I'd like to learn how to do this as well.
peab 21 hours ago [-]
Wow that's awesome
verdverm 1 days ago [-]
Seriously, I'm using a Qwen 3.8 27B on the homelab, distilled from supposed Fable traces. Regardless, the difference is notable, less thinking, better output. Distilled / heavy quant is better than the original (imv)
side quest, are fable distillations only wrong when it's another country?
soundworlds 22 hours ago [-]
See, you should now share it, so others can benefit without everyone having to do the same re-training :)
PEe9bB7D 1 days ago [-]
i also need more info!
GodelNumbering 1 days ago [-]
I am thinking about opensourcing everything, although this is not my main domain or my main startup, so the overhead of huggingface etc seems a bit unnecessary
Edit: will do as soon as possible
jack_pp 1 days ago [-]
just ask the agent to write it up if you don't have time to do a write-up yourself
genxy 23 hours ago [-]
Just use the post-one-off-project-to-huggingface-skill.md
equinumerous 1 days ago [-]
+1, would like to see. Even if it's not fully "ready for consumption", it's probably enough to reproduce the results.
atombender 1 days ago [-]
Would also love to read a write-up about this!
jjice 1 days ago [-]
Please do! Small, specialized models need more love and the time you spent would be a gift!
onel 14 hours ago [-]
Please do share it.
yashthakker 23 hours ago [-]
[dead]
okamiueru 1 days ago [-]
Golden age before the age that ends humanity. Not talking about any "rogue AI", just the known statistical models of what is coming due to climate change.
verdverm 1 days ago [-]
Do those statistical models account for declining birth rates or are they based on prior population growth projections?
19 hours ago [-]
xhevahir 1 days ago [-]
> what a time to be alive!
It's good to hear you're enjoying yourself, but I suggest retiring that expression. It's really beginning to grate.
willy_k 19 hours ago [-]
It’s not so good to hear your pessimism, but I suggest retiring spreading it online. It’s really beginning to grate.
JSR_FDED 19 hours ago [-]
Ehh, I’m more annoyed by people starting comments with “ehh”
tukHelix 1 days ago [-]
It’s the first time I know fireworks has a team doing model research. I do have a complex mood in that. On one hand, I’m always happy to see improvement of OSS models, whether that’s on intelligence or cost-efficiency. On the other hand, I would be a little worried about using fireworks as my API provider. Till the moment I saw this news, I had been using fireworks as my provider of deepseek v4 flash, because I thought fireworks acting as a role deploying OSS models and selling calculation resources, should be safe to use without worry of data being used for training since there’s no “conflict of interests”. But I would think twice now.
bradfa 1 days ago [-]
Just read the terms of service and read this blog post and I think your concern will be addressed.
this work may explain why recent models like qwen-3.8-flash and MiMo-2.6-* have not made it into their offering, which has given me reason to pause my excitement for Fireworks
noodletheworld 22 hours ago [-]
Seems irrelevant? Of course we don’t use data for training.
…trust me bro.
It’s obviously easier to believe when they’re not training models.
Eh, anyway this whole thing is just an ad:
> Looking to take Ember-1 one step further, and optimize it for your use case? We are also launching training support for Ember-1, enabling enterprises to build customized, token-efficient models tailored to their needs with their own data. The future of open models is specialized models trained on your specific workload.
Probably, I guess, fancy serverless infrastructure actually makes virtually no difference to hosting really large models that people want to use, and “just” being an inference provider for open weight models turns out to have no moat.
So this is a bit of a pivot to “use our training infrastructure too…!” imo.
Pivot? Sure. Go them. Not what I signed up for though. /shrug
mrngld 11 hours ago [-]
I think legal jurisdiction matters in discerning these things. The EU or US both have courts that, despite anyones opinion, are regarded as having robust contract enforcement. If Fireworks is domiciled in either then claiming ZDR and instead training on the data would be an enormous financial footgun. There's reasons why companies on both sides of contracts, even those with little to no US or EU activity, agree to use US or EU courts for enforcement.
Hong Kong I think used to be a popular option as well, before the handover from the UK.
Being terminally 'online' can make people cynical about everything, but have to temper things with reality a bit too.
vohk 5 hours ago [-]
Another factor in the cynicism is that robust enforcement is very 'pay to play'. I think your point is accurate for a company like Fireworks, largely doing business with peers that could afford litigation if necessary.
Good luck to an individual or small business trying to hold OpenAI or Meta to account if they breach contract over training, or do something crazy like copyright infringement on an industrial scale.
shostack 1 days ago [-]
Last I checked they still offer no training ZDR US based hosting. It is one of my three pinned providers for deepseek v4 flash along with Parasail and Deepinfra.
slim 1 days ago [-]
That also could explain why openrouter is worth that much
owl_might_ 1 days ago [-]
[flagged]
netvarun 1 days ago [-]
Off topic:With sol pricing drop tbh kimi k3’s value prop has not been that great. For our internal use case/testing/benchmarks sol come out with way better quality and much cheaper costs.
Kimi really needs to drop their pricing (I heard it’s set by them across all the neoclouds)
Sol is at 2/10 vs kimi’s 3/15
nicce 1 days ago [-]
Sol pricing dropped but so did the quality few days ago. I wonder when these companies are sued for making the terms from their side to go downwards while taking the same subscription cost.
solarkraft 1 days ago [-]
Is anybody tracking these quality changes? All I've seen so far are accusations (quite a few at this point) but not really any actual data.
Gareth321 15 hours ago [-]
They're impossible to objectively track by design and intent. Intelligence is such a nebulous target that one can find any number of metrics to support any premise: that the model is smarter, and that the model is dumber. We use benchmarks to attempt to standardise comparisons but these are quickly ingested into the training data and then become effectively useless. You might have heard the term "bench-maxed." Meaning that a benchmark has an effective lifespan in months.
They've been using the Pelican test in the /r/Codex subreddit with some success. One major finding is that OpenAI has been silently degrading the model while charging Astra prices. Another finding is that even when the model has not been silently degraded, pelican quality is significantly lower. Sometimes comically so. The general consensus right now is that the new GPT-6 Sol model is an updated Terra model. Many intelligence metrics are roughly similar. Meaning the most recent model updates were an attempt to rebalance compute rather than improve intelligence.
Ultimately I've never seen users as upset about GPT-6 Sol/Luna than I have right now. Even Astra has been noticeably degraded for me and everyone else I have asked. This is compounded by the fact that Opus 5.5 is a generational improvement at an affordable price. There is currently no competition.
vlyan 11 hours ago [-]
it's been known since the early days of GPT-4 that they alter the models whenever they feel like it.
Haven't looked into how accurate the page is, but the list of regressions on the bottom looks terrifying, at first glance?
copperx 22 hours ago [-]
Yes, it looks like regressions are frequent, but sometimes performance goes back to baseline quite fast.
bpavuk 1 days ago [-]
how the hell do we even track that? and before someone says...
—"Benchmarks!"
...I'll tell that they can be gamed so easily, and they are on a consistent basis.
solarkraft 1 days ago [-]
Sure they are, but do you think they are continuing training to improve a model after release without bumping the version number, presumably only to game the benchmarks?
nananana9 19 hours ago [-]
If you're willing to cheat, isn't it just a matter of grepping for the benchmark's question and pasting the solution in the chain of thought?
conorcleary 4 hours ago [-]
like windows 9
hn8726 22 hours ago [-]
100%, I wish for a legislation which would require the providers to give you at least a unique hash identifying the model (and infra running it, if it affects output) - such that the same hash must give the same output given the same seed. Right now it's all just vibes
koyote 1 days ago [-]
I am glad I am not the only one to notice. I feel like I've gone back to Sonnet 4 levels of incompetence!
With Sol 6 I am back in a world where the model writes bad code because it is lazy ("You're absolutely right, I did not [do it properly] because I did not want to edit [a normal amount of files]").
cromka 5 hours ago [-]
This is my experience today working with it:
- it actually failed to correctly understand a simple English grammar and logical implication of it, then when challenged it admitted its mistake but couldn't explain why it made it.
- for the code I am working on, I asked to create two PRs for the two small features (couple lines of code). It created one in upstream, as intended, and other one in my own fork. Just like that, out of nowhere, and called the job done.
- it said it would ask me to approve/ammend the suggested PR message, it never did and fired off right away
- it keeps forgetting the changes it did itself; no context compaction was used
- it said it tested the change visually, but it did not even try
On top of that, it ignores all of my AGENTS.md, which is short and concise. I mean I point it at ignoring it, it acknowledges and ignores again.
This is astonishingly bad and it is nowhere close to Sol 5.6, or even DeepSeek 4.1! I swear even Gemini 3.8 is slightly better.
To me, Sol6 is what Opus5 was for Claude.
I keep saying that with self-hosting, you at least know what to expect and don't have to trust they nerf their models as they go. I was a skeptic and considered nerfing a conspiracy theory, but at this point with enough experience, I have experienced enough to fully see this being a thing.
Probably only a matter of time before some class action happens.
1 days ago [-]
pornel 1 days ago [-]
Competition is good. Without K3/GLM/DS4 etc. there would be no pressure on OpenAI to drop Sol's price.
alansaber 13 hours ago [-]
Exactly, competition between both frontier model companies and the chinese labs is the primary factor suppressing consumer prices.
drob518 1 days ago [-]
Agreed. Even on the open weight side, GLM 5.3 has roughly equivalent performance to Kimi K3 for less than half the cost.
segmondy 10 hours ago [-]
No it doesn't. 5.3 is great, it's not K3 good.
drob518 4 hours ago [-]
Fine, let’s use your numbers, whatever they are. Either way, GLM 5.3 is a lot better than half as good as Kimi K3.
k__ 1 days ago [-]
With DeepSeek's pricing, no other value prop has been great.
conception 22 hours ago [-]
Mimo has entered the conversation.
nostrebored 1 days ago [-]
Agreed, I think the only place where it’s still interesting is ui design. Visually kimi and muse feel much nicer than frontier models to me, but maybe it’s an artifact of everything terrible being Claude Design
copperx 23 hours ago [-]
GLM 5.3 is great too. Thanks for suggesting Muse.
7777777phil 1 days ago [-]
I was surprised by that. I run my benchmark [1] every couple of days and was sure this model will be ath the pareto frontier, if not THE pareto frontier. But no:
Ember isn't picked yet. In planning, Opus 5.5 wins under the planning weights. In code, GPT-6 Sol dominates it: also 10/10, but with a higher quality score and a lower estimated cost. Ember has no intelligence index, so its starting score is only 0.73, which holds its 10/10 down to 0.954 against Sol's 0.975.
The more I learn about Fireworks the more unsavory they seem as a company. I don’t care what the license says, Moonshot has been openly improving, sharing research, and providing weights for the models that make up your entire bottom line, and the moment you can improve them in reciprocal it’s closed weights, “this is our own proprietary” nonsense? Where are we that China has better open source ethos than America?
peri-cl 1 days ago [-]
Why is proprietary-licensed software unethical?
Kimi K3 itself isn't FOSS. Speaking of reciprocity: Fireworks is presumably paying Moonshot serious money for the right to do what they are doing here, since Kimi's license[0] excludes commercial inference providers (such as Fireworks) from gratis use. It requires them to: "...enter into a separate agreement with Moonshot AI before using the Software or its derivative works..."
What you say is true, and Moonshot definitely has an agreement with Fireworks, but it still just feels wrong. It’s not the America I grew up in, where if you took, you gave back. I know that’s not a very coherent and practical position, but it feels true to me.
bonoboTP 12 hours ago [-]
They are giving back - money to Moonshot, most likely. Why should they give "back" to you? Did you make Kimi K3, so that "back" makes sense?
anentropic 11 hours ago [-]
is that really the America you grew up in...?
byzantinegene 20 hours ago [-]
presumably...
tkamado 3 hours ago [-]
in case people forgot: Fireworks with Cursor together rebranded K2.6 to Composer 2 without proper attribution
jamienk 1 days ago [-]
Ignoring for the moment issues of what "counts" as open, won't open models rapidly advance due to stuff like this in ways that it's less possible for the proprietary ones to do? This is exactly how Linux & Wikipedia, for example, overtook their "frontiers", right?
andsoitis 1 days ago [-]
> Ignoring for the moment issues of what "counts" as open, won't open models rapidly advance due to stuff like this in ways that it's less possible for the proprietary ones to do? This is exactly how Linux & Wikipedia, for example, overtook their "frontiers", right?
I suspect the advantage that catapulted Linux ahead of the establishment was less technical potential and talent and more organizational advantage. That's not to diminish the technical talent of the Linux crew, but them being unencumbered gave them more degrees of freedom. The rest is history.
So as long as the AI companies don't succumb to "big company" dynamics, they can outlead. To wit: Open AI and Anthropic are kicking Google's ass.
mirekrusin 1 days ago [-]
I think people make mistake here, google’s approach is not to spend $2.3 on every $1.0 earned, they’re riding on serving to masses “luna”, they absolutely have way more powerful models internally but they don’t clutter their infrastructure with fragile and costly intelligence-of-size inference frontier. I think “underdog” perception is illusory/temporary, not stupidity - calculated, conscious, longer term bet.
kingstnap 1 days ago [-]
Where do you get this "not to spend $2.3 on ever $1.0 earned" from?
Google might not have compelling frontier offerings, their chat harness is complete garbage compared to any other lab (in large part due to a bizarrely badly designed harness where something like code execution requires the prompt to undergo some sort of classification step, no idea what they are doing).
But they absolutely kill in terms of usage offerings. Google lets one subscription be used by *SIX* different google accounts on a family plan.
Plus I currently literally get *$40/month* of Gemini API credits on developer.google.com because they gave me a $10/month grant 4 times.
They give you 200 cloud compute units on google collab, this literally lets you spin up an H100 for around 40 hrs or something if you want to try spinning up local models.
You get Jules (huge allotment btw), Image gen, Video gen, Music Gen, antigravity usage, 5 TB of cloud storage, Notebook LLM...
hn8726 22 hours ago [-]
Okay 5tb cloud storage is their most expensive plan. But what do you actually use video or image gen for? Or music gen? Antigravity is garbage, I guess you can do some stuff with Gemini models over API.
I was paying for Google ai and then realized that between obscure limits and gimmick features I don't really need it. Canceled my subscription and didn't even notice a difference
mirekrusin 18 hours ago [-]
It's an example number, which doesn't matter much but it comes from widely reported $2.30 spent on ops for every $1 of revenue in spring of 2025 for Anthropic.
bpavuk 1 days ago [-]
I tend to agree, but I also should highlight how expensive this shit really is.
in one month, Google actually went cash-negative. [0] even still, they are subsidizing their stuff a lot less, have the most opaque and variable limits, and increase adoption through bundling and shuffling features. I can't even share my Google One storage without subscribing to a Google AI plan anymore, but previously any plan except Google One Lite was shareable.
if you tell me that's not enough to go after frontier, then how much money are Anthropic and OpenAI burning?
Diff people have diff motives to experiment, then new work is done on top of stuff that "hits" in a way no one anticipated. Then work gets piled on top in a way that might make it hard to port
andsoitis 1 days ago [-]
> then new work is done on top of stuff that "hits" in a way no one anticipated.
Indeed. And when you have freedom to play, you are able to find new stepping stones that you didn't anticipate. And you can combine stepping stones in new ways to make new discoveries.
Greatness cannot be planned.
segmondy 1 days ago [-]
No, because close labs/models borrow but don't contribute back.
Won't the "frontier" labs figure out whatever techniques were used and apply them to their closed models?
k__ 1 days ago [-]
If they can keep up.
The lock-in is less pronounced as it is with AWS or MS.
cyanydeez 1 days ago [-]
Like how the last 2 decades of tech companies are thinly veiled open source pilfering into business units.
sincerely 1 days ago [-]
"Oh darn, you know that thing I made and released with explicit, precise language defining who can use it and what, if any, restrictions apply? Well now someone is using it in complete accordance with those conditions I set out, and that's somehow making me upset"
zeroq 1 days ago [-]
The difference between contributing to OS and AI, is that the first is a hobby alternative to woodworking or hiking, while the other can easily bootstrap you a company you can get millions in investment, at least for time being.
intothemild 1 days ago [-]
Yes, absolutely, but only if people keep contributing in the open.
jack_pp 1 days ago [-]
not necessarily, just knowing something is possible will motivate others to achieve it somehow. Which is why there are so many LLMs and OAI doesn't have a monopoly
Arcuru 1 days ago [-]
Over on /r/LocalLLaMA there's a group that's been getting popular doing the same thing for the Qwen 27B (and other) models. - https://huggingface.co/ukisai
alentred 15 hours ago [-]
Does anyone have a real-world usage feedback on this model? Because the benchmarks look, I would say, too good: it scores even higher than Qwen 3.8 27B in some benchmarks - I am not even sure how is it possible. So the question is if this is benchmaxxing or is it actually good for practical usage like coding in complex projects, research, etc.
tangled 1 days ago [-]
What am I missing here? I think of fireworks as an inference provider serving open weights model. The value that they primarily provide to customers is that (i) they improve reliability by balancing across a bunch of clouds/neoclouds, (ii) they get better pricing by buying capacity in bulk, and (iii) they reduce operational costs. So far so good.
I can also see the argument for providing a post-training service from a customer acquisition perspective: "hey, we can fine-tune this open weights model, so it both gives better/more predictable results than OpenAI/Anthropic and also is cheaper. And btw, once we've won your business, please run this model on our infra."
But what I'm struggling to understand is fireworks spending a bunch of money (on salaries and compute) releasing a frontier model that is going to rapidly fall behind the frontier. Is this "just" advertising for them, both for customers and also for hiring? Or are they actually trying to stay on the frontier? If so, to what end?
danielmarkbruce 1 days ago [-]
The end: make lots of money.
The means: systematically take existing reasoning models, do some more post-training of some sort to make them achieve the same outputs with less reasoning tokens (ie, cheaper). Same quality but cheaper is always valuable.
It's unclear if they can do this systematically and it's unclear if they can do it better than others. But, lots of things are unclear in AI at the moment, this doesn't seem outrageous on the surface. And, it could just be marketing. And it could be the first option with the backup of the second.
criemen 1 days ago [-]
> to what end?
I'd expect that their business strategy is to compete in more markets, and if successful, they can capture more value. This is the "easiest" for them as they already have GPUs, a training environment etc. For that platform it's not the worst if there's an internal customer team that can help shape the future and provide immediate feedback, and if it results in a good model, even better.
Other things I'd not be surprised they offer in the future in the same vein: A multi-model harness, coding agent (cloud and local), and maybe at a later point in time even a CPU-only cloud compute product.
TomasEkeli 1 days ago [-]
I think they are trying to show potential customers what is possible.
conception 22 hours ago [-]
Why does Cursor or Devin make their own models? If you use their models, then you can't fallback to other people's models on openrouter or anyplace else. You just stick with them.
k__ 1 days ago [-]
Vertical integration.
soerxpso 1 days ago [-]
I don't see why I would be interested in this model, considering the price difference. They advertise that it's the same as Kimi K3 in half the tokens. But the pricing is double the pricing of K3. So why do I care if it uses fewer tokens, if I'm paying double per token?
bigmadshoe 1 days ago [-]
Because you care about how many tokens are used per task. What you said is like only caring about the price of gas and not gas mileage of your car.
verdverm 1 days ago [-]
it's not exactly the same, the model stills "weighs" the same
here, it does less work, it's more like driving half as far but still paying the same total cost
this being said, K3 and E1 models are priced the same at $3.00 / $0.30 / $15.00
From reading the blog post, it is essentially exactly the same as the car example. It delivers the same performance on tasks, but using 40% fewer tokens. This is the same as a car getting you to the same destination but wasting less energy on excess heat, wind resistance, or whatever else affects fuel economy (I am not an expert, obviously). I am paying for an LLM to complete tasks for me, not for the intermediate tokens.
verdverm 24 hours ago [-]
I'm with you, I wrote prior comment under the assumption that they were priced differently (from GP claim as such), but they are priced the same (on Fireworks)
Will be taking Ember-1 for a spin on Monday and hopefully enjoy those better MPGs
ranguna 14 hours ago [-]
Why are you saying it drives half as far when both models complete the same task but one uses half the tokens?
Wouldn't the op be more correct with their gas/distance comparison?
Because both cars get to the same end destination (complete the same task).
Unless your end goal is to see the token numbers go up, but I'm not sure why that would be of interest.
zupa-hu 13 hours ago [-]
Agentic work costs ~input^2. For a single message, they cost the same. For long conversations, you end up paying much less.
seizethecheese 23 hours ago [-]
Half the tokens presumably means tasks get done twice as fast.
XCSme 12 hours ago [-]
I mean, Gpt-6 Luna seems a lot better in every aspect:
K3 TOS says they must sell no lower than what Moonshot charges. If you have a provider selling for less than $15/mtok, they are violating TOS from Moonshot.
however the listing on open router has a `/fp4` suffix, so perhaps this is an unlisted, quanted model for a lower price?
verdverm 21 hours ago [-]
as the saying goes, you get what you pay for
we require ZDR and Fireworks provides that on contract, so for us they are the same price
indigodaddy 20 hours ago [-]
Check out Neuralwatt. They are ZDR and great energy based K3 pricing. They also have a K3-fast which is basically no reasoning (in addition to regular K3).
verdverm 20 hours ago [-]
that is nothing like how we consume Ai
industry standard is price-per-1M tokens, don't do something different, even Google caved and moved from their char based pricing to tokens (the fundamental unit of computation in ai)
GPUs are rented in $/h, like every other piece of hardware in cloud
Bolwin 17 hours ago [-]
Why are you assuming the gpus are rented?
Anyway, they still have token based pricing if you prefer. The energy pricing is often cheaper though
This is partly the appeal of Jev et al; having a quick model for simple tasks, that doesn’t require that much thinking
It’s amazing all the workflows that models like that can unlock. And yes, classifiers and other ML models have been around for a while for these types of tasks, but Jev has made it easy and cheap to play and experiment. This in turn, is incentivizing people to try them for a bunch of stuff, unlocking creativity and producing a lot of new cool (and eventually potentially very useful) applications
demibabs 1 days ago [-]
What are the useful applications of Jev so far? Not to sound dismissive, I just haven’t seen what people are using it for yet.
vulture916 1 days ago [-]
Here's a third-party (not Jev) showcase of things people built, which helped me kind of get the appeal. https://bentossell.com/jev/ (not mine).
grosswait 24 hours ago [-]
LLMs can be too creative and often too verbose. Sometimes there is a right answer and a way to get there with the understanding of language, but despite using structured outputs, the model insists on inventing variations not in the schema or coming up with something completely different. A model like Jev that can not do those things, and can give the same output every time with given the same input, and be able to measure probabilities has many use cases.
neosat 1 days ago [-]
Lots of use cases!
I've personally used it for the following:
1. Evals (once you have your rubric defined and tuned using a reasoning model, jev can be great for running periodic evals especially those that run daily.
2. e-commerce catalog classification
3. quick search using anything as context and query mapping to a pre-defined set.
computerex 1 days ago [-]
At least for 1, evils, you’d want to use a good old reasoning model to get the best eval results.
elcomet 1 days ago [-]
Why not using a cheap LLM with thinking completely disabled ? I don't think it will be much more expensive than jev.
nico 1 days ago [-]
I’ve tested this with some local LLMs and their accuracy is in general better than Jev/Laya, but they are super slow in comparison as well
For example, a typical/stock LLM can’t really play Doom in real time, but a Jev-like model can. Just because of latency
Of course, if you want the best Doom player, there are way better and faster adhoc models
ssivark 1 days ago [-]
LLM inference has two very different regimes of work: prefill & decode. You can think of the former roughly as processing a pre-specified prompt, and the latter as sequential processing (auto-regressive token generation) eg. "chain of thought". The latter is very important for LLMs and cannot be ignored; it deeply influences infra design, even necessitates copious amounts of high-bandwidth memory. Jev-like models can ignore the latter and therefore optimize much better for the former, consequently operating at both better cost and latency.
andsoitis 1 days ago [-]
> The problem: thinking models think too much
Analysis paralysis stifles not just human intelligence, but other intelligences too.
condour75 6 hours ago [-]
Hamlet's soliloquy but it's about the number of b's in strawberry
minimaxir 1 days ago [-]
The thinking traces on some Chinese models just output the full response in the thinking trace, then output it again to the user, which is redundant.
AraneaDev 1 days ago [-]
Yes and thar makes you wonder if the Paradox of Choice would apply as well ;)
The more options you have, the harder it becomes to be satisfied with the one you picked.
dijit 1 days ago [-]
Given what an experience i had with Ember-2… I’m not sure I’d want to engage with its predecessor.
The pareto frontier needs clearer distinction. Benchmarks miss half the story. What, if any, capability is lost by the token reduction (for example, was it like super awesome at Golang before and now kind of sucks? that kind of distinction).
drob518 1 days ago [-]
Unfortunately, it’s hard to make a chart of that.
ttmoab 1 days ago [-]
[flagged]
intothemild 1 days ago [-]
So they trained a model on open weights, and then aren't releasing the weights... am I reading this right?
netvarun 1 days ago [-]
Technically kimi k-3 weights license is not open weight (it has a lot of restrictions). I would classify it as ‘weight open’ similar to the bsl and fsl ’source open’ licenses.
Evidlo 1 days ago [-]
weight available
DonsDiscountGas 1 days ago [-]
It happens. Most open licenses aren't GPL style copyleft.
kingstnap 1 days ago [-]
There is little to no point reading the article as well. It's stripped of all alpha.
> task and environment feedback
> on-policy planning and learning
> feedback connects decisions to their consequences
These are deliberately the least informative phrases you could possibly use to describe what you have done, while still being in the realm of words that go over a generic investor who has no idea whats going on and may be dazzled by sciencey sounding language.
Cursor compose 2.5 article where they used and described on policy self distilation was actual alpha.
intothemild 1 days ago [-]
This is precisely my point.
makeramen 1 days ago [-]
Aren't Cursor Composer models like this too? At some point all the extra RL you do can be considered as proprietary information added.
Not suggesting this is right or wrong, but is sort of the nature of the technology.
1 days ago [-]
swagatkonchada 1 days ago [-]
It happens with open source software all the time, why would we expect any different with open source weights.
reactordev 1 days ago [-]
Because we do. The GPL isn't a suggestion. If you can take open source code and make private software out of it then what are we all doing? No, license requirements and agreement are law for a reason.
Cut my teeth on Ember back in the 1.x days. The learning curve was steep, but the productivity once you got it was immense.
owenthejumper 7 hours ago [-]
I don't get it. Half the tokens, but triple the cost on Open Router? So I pay more? Great idea otherwise...
k__ 6 hours ago [-]
Half the tokens can also mean quicker task completion
spdustin 1 days ago [-]
Been thinking about the feasibility of training a model using synthetic thinking traces that were reduced to caveman-speak prior to being used for training. Seems like it would be fairly easy to generate plenty of suitably lobotomized synthetic traces with a pair of cheap-ish models. Or even just using good old fashioned NLP to aggressively remove stop words and reduce trace words to lemmas.
fbrncci 23 hours ago [-]
I am no longer going to be impressed by new model releases unless they introduce an entirely new paradigm of interacting with them, that is going to make the benchmarks look like everything else isn't even 5% as capable.
madisonkanna 23 hours ago [-]
I work at Fireworks and it's cool to see this was posted.
I'd be interested to hear what people found most interesting about Ember, and what kinds of follow-up research or educational material would be useful to you all?
hankbond 22 hours ago [-]
What made you choose/advertise using Doximity's benchmark? I use to work there and it was interesting seeing it pop up.
olgava 14 hours ago [-]
[dead]
qeternity 1 days ago [-]
This is undoubtedly great. But most of the inference cost today for dominant use cases (agentic coding) are in the prefill, not the decode. This is one of the reasons that DeepSeek is so aggressively optimizing prefill and caching.
gitghxst 12 hours ago [-]
that's really interesting. would people be interesting in reviewing it at https://unbenchmark.com?
srameshc 1 days ago [-]
> The problem: thinking models think too much
I see that with Opus 5, it started thinking like crazy in the last few days , I don't think my workflow is that complicated, still it gets into thinking mode and stays there
Need this done for DeepSeek, ideally one of the Flash models.
drob518 1 days ago [-]
And GLM. Both Deepseek 4.1 Flash and GLM 5.3 Flash are quote verbose when thinking.
atemerev 1 days ago [-]
If you have the compute, I have the expertise.
riquito 1 days ago [-]
Aside. I find the "cost per task" charts both useful and uncanny. Is It better a model that takes me to 90% in 1 dollar or one that takes me to 95% in 2 dollars? Or a different model that too scores 90% in 1 dollar? How much will it cost me the last 10% or 5%? At the end of the day, cost to 100% is what matters and the half (90%) backed solution may require more to reach 100% (or not, who knows?)
swiftcoder 1 days ago [-]
> Is It better a model that takes me to 90% in 1 dollar or one that takes me to 95% in 2 dollars?
It's pretty important to understand if your own work domain is one where the last 5% matters. In a lot of day-to-day software engineering tasks, it doesn't, and one can get crazy mileage out of the cheaper models. OTOH, if you are performing novel research, that last 5% may be worth whatever it costs...
entrope 1 days ago [-]
The 90% and 95% are against some blend of tasks meant to be broadly representative. A pricey model seldom fails a problem that cheap models do well, so there's stratification of tasks by difficulty. Someone doing novel research may be in the "hard" 15% of the blend, where P(solution) goes from one third to two thirds.
On the other hand, if it's cheap to tell whether you got a good solution, and you think the 90 and 95% apply to your task blend, then it's almost always worth trying the cheap model first.
jonplackett 1 days ago [-]
Could be a really interesting article but they disabled reader mode so I guess I’ll never know.
ed_mercer 20 hours ago [-]
So are they going to release this tuned model? Or keep it for themselves?
pullstart 14 hours ago [-]
Still remember diving into Ember back around 2015. Always appreciated its strong conventions, made for really maintainable apps in the long run.
sailfast 23 hours ago [-]
Is this a useful model or just an ad for Fireworks runtime? I can’t really tell…
tdhz77 1 days ago [-]
Does anybody know if this would be a good model for creative writing?
combobyte 1 days ago [-]
> model
> creative
Choose one.
tdhz77 1 days ago [-]
Are you a bot?
combobyte 1 days ago [-]
You really have no sense of irony, do you?
dmkolobov 1 days ago [-]
This is cool! But also: am I wrong for thinking “Pareto frontier” is some pretty silly/clever marketing jargon? Is this common phrasing for basically saying: test performance per spend on tokens is decent?
sebzim4500 1 days ago [-]
I don't see why? It is a well defined term that existed prior to the recent AI bubble/revolution, and from what I can see they are using it appropriately.
dmkolobov 1 days ago [-]
Fair enough!
blissofbeing 1 days ago [-]
Would be nice to include in fire pass.
dbuxton 1 days ago [-]
Do they mean Opus 5.5 or Opus 5?
ls612 1 days ago [-]
On the smaller end, Quen 3.8, while being extraordinarily capable for a small local model, also suffers from extreme thinking. I wonder if the techniques described here generalize to other models too.
spijdar 1 days ago [-]
I suspect it might generalize to other large models, but I don't think Qwen3.8 27B is one of them. Kimi K3 is a 2.8 trillion parameter model, and I suspect that is playing a big role in being able to reduce the length of CoT without taking a hit in quality.
I don't think the article mentions Pareto frontier enough.
Also, did I miss a memo? Suddenly every article on AI seems to be talking about the Pareto frontier - or have I just not been paying attention?
DonsDiscountGas 1 days ago [-]
They want it to be the best at something. And it's obviously not the absolute smartest. So here we are.
AnodicElegy 1 days ago [-]
I guess they figure "best bang for your buck" comes off a little too colloquial.
swiftcoder 1 days ago [-]
I would really love if we brought back some colloquialisms in this field. Not that long ago most folks in tech would have had pretty blank looks on their faces when someone started talking about the "Pareto frontier"
user43928 1 days ago [-]
Pareto frontier on some benchmark that I am hearing of for the first time.
Kimi K3 with less reasoning tokens isn't exactly exciting either, and particularly so if the license is less open than original Kimi K3.
intothemild 1 days ago [-]
Is there a Pareto frontier for the number of times articles mention or don't mention a Pareto frontier.
alienbaby 1 days ago [-]
when everyones fighting to be 'somewhere in the pile' they need some way to advertise they have made progress while not being the best.
gradous 23 hours ago [-]
code
desireco42 24 hours ago [-]
I am confused over this... I get efficiency but price seems too high to me. I didn't try the model, so it migth be beyond fast or some other quality that is not obvious.
Anyone knows more or used this model?
themgt 1 days ago [-]
The result? Ember-1 set a new Pareto frontier for Bedside Bench across both open and closed models including GPT-5.6 Sol, GPT-6 Astra, and Claude Opus 5 on cost/task.
"Pareto": 8 hits
"Opus 5.5": zero hits
wmf 1 days ago [-]
Obviously this research was done before 6.0 Sol and Opus 5.5 came out. Your point stands that the frontier moves quickly and small gains can be eclipsed quickly.
logicallee 1 days ago [-]
This is really interesting. I think the Fireworks Serverless Training infrastructure they used to develop it is also unique and needed. Except if someone works at one of a handful of the largest labs, it is very difficult to set up or try any sort of training pipeline. The managed training infrastructure makes it available to more people.
nostrebored 1 days ago [-]
I can’t help but think it’s more expensive tinker.
esafak 1 days ago [-]
It looks like it would be similar to GLM 5.3 Flash, had they tested it...
A tiny, smaller than 1b parameter model, fine-tuned, can kick ass for constrained work. I do not have a lot of budget, I fine-tune only on a 16GB M4 Mac Mini. But that also tells me the potential is wild. Progress has been slow since I moonlight on this.
I have been trying to build a set of models + agents for full-stack development, where each model does only a small piece, like take user prompt and break into backend/frontend tasks. Then a Rust+Diesel model, a Rust+Auxum model, a Solid+Router model and so on. I know this is wild but this is just theory - can 5 or 6 Qwen 3.5 0.8b models do full-stack web development? My hunch says they can, better than what most people expect. Heck, with a good harness, it might beat all the cheaper models for the specific task, like Haiku or Luna.
That being said it’s much easier at the moment to continue to use the frontier providers for most general tasks, that is the argument I’ve heard.
For creating these types of fine tuned local models, on constrained hardware for inference, I do think this is the way to go for specific tasks too!
Have you actually fine-tuned yourself? Email categorization comes to mind and there are tons of non-LLM approaches even that will give fantastic results. How did spam filters work before LLM?
I think LLMs just made us think that is the only way. It is not.
Or are the subagents generating your training data using a closed/paid model?
For everyday work that happens frequently it's better to have a tiny specialized model instead of making billable API calls or turning your laptop into an 80W space heater for 20 seconds to run a general purpose model.
The large models can be used to generate synthetic training data. Tell them to make up 100,000 tasks paired with the resulting output as a 1-time cost. Then use that to train a small model.
Think of it as distillation, but focused on a specific task.
Which 5$ is a pretty easy sell if its useful in any way, It's pretty easy to justify a purchase if its yours forever and doesn't use much CPU so is easy to run I mean people were spending 1000$+ on mac mini setups to run local llms or run remote agents.
They also used Astra for the coding, which they can't get on a $10 subscription. And then there's the actual training cost.
This particular example is maybe a niche, but 1400 people can use a few hundred queries in a reasonable amount of time.
I think OP's point remains, if you generate 140k pairs, your local model would need to run that many to offset having just used the generator (SOTA or not) model to begin with.
I wonder if another approach if latency is a concern is just to do a two shot pass with Jev (perhaps given small context you'd want one to match command, then one to match args of given command) would be an extremely fast, and cheap way to do it - rather than training your own.
Not quite. One, because you save on the initial query being sent multiple times. Two, because the reasoning will be very similar at the beginning; "user asked me to", "let's check what's in this project already", etc. you'll get similar actual output cost, but input and reasoning will be shorter.
I have a single command that fires up llama.cpp on cpu only using gemma4 e2b, answers a single question from the command line and exits. This takes about 3 seconds to load from an SSD, and is smart enough to solve exactly these "remind me of the syntax" scenarios if you dont wanna switch to a browser.
When I do something "heavy", my 9800x3D will kick on the fans and start making lot of heat. This is fine when I'm intentionally doing say a transcode, but if I'm just querying syntax it could get pretty annoying. Do you "feel" it? I know that will be pretty machine dependent.
I ask cause would this be a kind of model distillation?
I have a small model I'm looking to train on some data, and I have some real live data but I'd love to be able to extend it.
edit: others have asked any you have replied "soon (tm)", looking forward for that day
Not my project
I’m really pushing my recall, but I want to say it was written in Ruby and stored pre-configured commands that it just did traditional search over.
I vaguely recall it working okay because 99.99% of the questions people asked were the same (“tar command to gzip a directory and strip the prefix” is something I google like once a month).
I am building a natural language to CSV/Excel commands for a "wrangler" type desktop app. Same issues. The MVP is being built with parsers of sorts, entirely code generated. Then I want to fine-tune a tiny model at some point.
https://github.com/brainless/baho
we dont know what the result is and how its impressive.
Aside on the aside, I welcome this new era of really personal software. Not Ai's being sycophants, rather being able to easily and quickly change, adapt, or extend software I am not familiar with.
This is the part where the narrator looks at the camera and says "Don't try this at home, kids!"
[Search: Can I refund Google cloud?]
It looks like we’re not able to ask for a refund since we did actually use all of that compute intentionally.
Would you like me to write you a pleading email to send to the support team?
Are you confusing this with an OAuth token or something?
https://docs.cloud.google.com/billing/docs/how-to/budgets-sp...
I had to intervene a few times. For instance, as smart as the models are said to be (Astra), it would copy the full training run, train on the server, pull every checkpoint to the local machine, then run tests, update. So, the bandwidth bill was as high as training bill for the first 6 hours. It could have simply tested each checkpoint on the server, saved time and money, didn't occur to it until I said.
I wouldn't put my house on it. Brave.
Can you (or your agent) please write a tutorial or share some good links.
(Namely the fine tuning part.)
I'd like to learn how to do this as well.
https://huggingface.co/vwdubb/Qwen3.8-27B-Fable-Distill-NVFP...
side quest, are fable distillations only wrong when it's another country?
Edit: will do as soon as possible
It's good to hear you're enjoying yourself, but I suggest retiring that expression. It's really beginning to grate.
this is our preferred open weight token vendor
this work may explain why recent models like qwen-3.8-flash and MiMo-2.6-* have not made it into their offering, which has given me reason to pause my excitement for Fireworks
…trust me bro.
It’s obviously easier to believe when they’re not training models.
Eh, anyway this whole thing is just an ad:
> Looking to take Ember-1 one step further, and optimize it for your use case? We are also launching training support for Ember-1, enabling enterprises to build customized, token-efficient models tailored to their needs with their own data. The future of open models is specialized models trained on your specific workload.
Probably, I guess, fancy serverless infrastructure actually makes virtually no difference to hosting really large models that people want to use, and “just” being an inference provider for open weight models turns out to have no moat.
So this is a bit of a pivot to “use our training infrastructure too…!” imo.
Pivot? Sure. Go them. Not what I signed up for though. /shrug
Hong Kong I think used to be a popular option as well, before the handover from the UK.
Being terminally 'online' can make people cynical about everything, but have to temper things with reality a bit too.
Good luck to an individual or small business trying to hold OpenAI or Meta to account if they breach contract over training, or do something crazy like copyright infringement on an industrial scale.
They've been using the Pelican test in the /r/Codex subreddit with some success. One major finding is that OpenAI has been silently degrading the model while charging Astra prices. Another finding is that even when the model has not been silently degraded, pelican quality is significantly lower. Sometimes comically so. The general consensus right now is that the new GPT-6 Sol model is an updated Terra model. Many intelligence metrics are roughly similar. Meaning the most recent model updates were an attempt to rebalance compute rather than improve intelligence.
Ultimately I've never seen users as upset about GPT-6 Sol/Luna than I have right now. Even Astra has been noticeably degraded for me and everyone else I have asked. This is compounded by the fact that Opus 5.5 is a generational improvement at an affordable price. There is currently no competition.
https://arxiv.org/pdf/2307.09009
the accusations are quite a few because people notice.
—"Benchmarks!"
...I'll tell that they can be gamed so easily, and they are on a consistent basis.
With Sol 6 I am back in a world where the model writes bad code because it is lazy ("You're absolutely right, I did not [do it properly] because I did not want to edit [a normal amount of files]").
- it actually failed to correctly understand a simple English grammar and logical implication of it, then when challenged it admitted its mistake but couldn't explain why it made it.
- for the code I am working on, I asked to create two PRs for the two small features (couple lines of code). It created one in upstream, as intended, and other one in my own fork. Just like that, out of nowhere, and called the job done.
- it said it would ask me to approve/ammend the suggested PR message, it never did and fired off right away
- it keeps forgetting the changes it did itself; no context compaction was used
- it said it tested the change visually, but it did not even try
On top of that, it ignores all of my AGENTS.md, which is short and concise. I mean I point it at ignoring it, it acknowledges and ignores again.
This is astonishingly bad and it is nowhere close to Sol 5.6, or even DeepSeek 4.1! I swear even Gemini 3.8 is slightly better.
To me, Sol6 is what Opus5 was for Claude.
I keep saying that with self-hosting, you at least know what to expect and don't have to trust they nerf their models as they go. I was a skeptic and considered nerfing a conspiracy theory, but at this point with enough experience, I have experienced enough to fully see this being a thing.
Probably only a matter of time before some class action happens.
Ember isn't picked yet. In planning, Opus 5.5 wins under the planning weights. In code, GPT-6 Sol dominates it: also 10/10, but with a higher quality score and a lower estimated cost. Ember has no intelligence index, so its starting score is only 0.73, which holds its 10/10 down to 0.954 against Sol's 0.975.
[1] https://philippdubach.com/posts/jev-model-router-for-pi/
6 or 5.6? Because 6 is hot garbage
Kimi K3 itself isn't FOSS. Speaking of reciprocity: Fireworks is presumably paying Moonshot serious money for the right to do what they are doing here, since Kimi's license[0] excludes commercial inference providers (such as Fireworks) from gratis use. It requires them to: "...enter into a separate agreement with Moonshot AI before using the Software or its derivative works..."
[0] https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE#...
I suspect the advantage that catapulted Linux ahead of the establishment was less technical potential and talent and more organizational advantage. That's not to diminish the technical talent of the Linux crew, but them being unencumbered gave them more degrees of freedom. The rest is history.
So as long as the AI companies don't succumb to "big company" dynamics, they can outlead. To wit: Open AI and Anthropic are kicking Google's ass.
Google might not have compelling frontier offerings, their chat harness is complete garbage compared to any other lab (in large part due to a bizarrely badly designed harness where something like code execution requires the prompt to undergo some sort of classification step, no idea what they are doing).
But they absolutely kill in terms of usage offerings. Google lets one subscription be used by *SIX* different google accounts on a family plan.
Plus I currently literally get *$40/month* of Gemini API credits on developer.google.com because they gave me a $10/month grant 4 times.
They give you 200 cloud compute units on google collab, this literally lets you spin up an H100 for around 40 hrs or something if you want to try spinning up local models.
You get Jules (huge allotment btw), Image gen, Video gen, Music Gen, antigravity usage, 5 TB of cloud storage, Notebook LLM...
in one month, Google actually went cash-negative. [0] even still, they are subsidizing their stuff a lot less, have the most opaque and variable limits, and increase adoption through bundling and shuffling features. I can't even share my Google One storage without subscribing to a Google AI plan anymore, but previously any plan except Google One Lite was shareable.
if you tell me that's not enough to go after frontier, then how much money are Anthropic and OpenAI burning?
[0]: https://www.techspot.com/news/113214-google-records-first-ne...
Indeed. And when you have freedom to play, you are able to find new stepping stones that you didn't anticipate. And you can combine stepping stones in new ways to make new discoveries.
Greatness cannot be planned.
The lock-in is less pronounced as it is with AWS or MS.
I can also see the argument for providing a post-training service from a customer acquisition perspective: "hey, we can fine-tune this open weights model, so it both gives better/more predictable results than OpenAI/Anthropic and also is cheaper. And btw, once we've won your business, please run this model on our infra."
But what I'm struggling to understand is fireworks spending a bunch of money (on salaries and compute) releasing a frontier model that is going to rapidly fall behind the frontier. Is this "just" advertising for them, both for customers and also for hiring? Or are they actually trying to stay on the frontier? If so, to what end?
It's unclear if they can do this systematically and it's unclear if they can do it better than others. But, lots of things are unclear in AI at the moment, this doesn't seem outrageous on the surface. And, it could just be marketing. And it could be the first option with the backup of the second.
I'd expect that their business strategy is to compete in more markets, and if successful, they can capture more value. This is the "easiest" for them as they already have GPUs, a training environment etc. For that platform it's not the worst if there's an internal customer team that can help shape the future and provide immediate feedback, and if it results in a good model, even better.
Other things I'd not be surprised they offer in the future in the same vein: A multi-model harness, coding agent (cloud and local), and maybe at a later point in time even a CPU-only cloud compute product.
here, it does less work, it's more like driving half as far but still paying the same total cost
this being said, K3 and E1 models are priced the same at $3.00 / $0.30 / $15.00
https://fireworks.ai/models/fireworks/kimi-k3
https://fireworks.ai/models/fireworks/ember-1
Will be taking Ember-1 for a spin on Monday and hopefully enjoy those better MPGs
Wouldn't the op be more correct with their gas/distance comparison?
Because both cars get to the same end destination (complete the same task).
Unless your end goal is to see the token numbers go up, but I'm not sure why that would be of interest.
https://aibenchy.com/compare/fireworks-ember-1-high/openai-g...
https://openrouter.ai/moonshotai/kimi-k3
however the listing on open router has a `/fp4` suffix, so perhaps this is an unlisted, quanted model for a lower price?
we require ZDR and Fireworks provides that on contract, so for us they are the same price
industry standard is price-per-1M tokens, don't do something different, even Google caved and moved from their char based pricing to tokens (the fundamental unit of computation in ai)
GPUs are rented in $/h, like every other piece of hardware in cloud
Anyway, they still have token based pricing if you prefer. The energy pricing is often cheaper though
A trust page is required like https://trust.fireworks.ai
Very happy with my $10/month OpenCode Go sub for personal use
https://trust.opencode.ai/
This is partly the appeal of Jev et al; having a quick model for simple tasks, that doesn’t require that much thinking
It’s amazing all the workflows that models like that can unlock. And yes, classifiers and other ML models have been around for a while for these types of tasks, but Jev has made it easy and cheap to play and experiment. This in turn, is incentivizing people to try them for a bunch of stuff, unlocking creativity and producing a lot of new cool (and eventually potentially very useful) applications
1. Evals (once you have your rubric defined and tuned using a reasoning model, jev can be great for running periodic evals especially those that run daily.
2. e-commerce catalog classification 3. quick search using anything as context and query mapping to a pre-defined set.
For example, a typical/stock LLM can’t really play Doom in real time, but a Jev-like model can. Just because of latency
Of course, if you want the best Doom player, there are way better and faster adhoc models
Analysis paralysis stifles not just human intelligence, but other intelligences too.
The more options you have, the harder it becomes to be satisfied with the one you picked.
https://en.wikipedia.org/wiki/Exapunks?wprov=sfti1
The pareto frontier needs clearer distinction. Benchmarks miss half the story. What, if any, capability is lost by the token reduction (for example, was it like super awesome at Golang before and now kind of sucks? that kind of distinction).
> task and environment feedback
> on-policy planning and learning
> feedback connects decisions to their consequences
These are deliberately the least informative phrases you could possibly use to describe what you have done, while still being in the realm of words that go over a generic investor who has no idea whats going on and may be dazzled by sciencey sounding language.
Cursor compose 2.5 article where they used and described on policy self distilation was actual alpha.
Not suggesting this is right or wrong, but is sort of the nature of the technology.
The words of a license are what the license is.
Looking at the hamsters drawings in this comparison, never saw any other model make such a similar version:
https://aibenchy.com/compare/fireworks-ember-1-low/google-ge...
I'd be interested to hear what people found most interesting about Ember, and what kinds of follow-up research or educational material would be useful to you all?
I see that with Opus 5, it started thinking like crazy in the last few days , I don't think my workflow is that complicated, still it gets into thinking mode and stays there
It's pretty important to understand if your own work domain is one where the last 5% matters. In a lot of day-to-day software engineering tasks, it doesn't, and one can get crazy mileage out of the cheaper models. OTOH, if you are performing novel research, that last 5% may be worth whatever it costs...
On the other hand, if it's cheap to tell whether you got a good solution, and you think the 90 and 95% apply to your task blend, then it's almost always worth trying the cheap model first.
> creative
Choose one.
That's just vibes, though.
Also, did I miss a memo? Suddenly every article on AI seems to be talking about the Pareto frontier - or have I just not been paying attention?
Kimi K3 with less reasoning tokens isn't exactly exciting either, and particularly so if the license is less open than original Kimi K3.
Anyone knows more or used this model?
"Pareto": 8 hits
"Opus 5.5": zero hits