Hacker Newsnew | past | comments | ask | show | jobs | submit | Alpha3031's commentslogin

Flip both, Cario Vmodei.

You mean complete control for open models, because you can run it on your own choice of hardware?

Enc-decs are usually harder to train at frontier scale. Not 100% sure what DeepSeek has done differently here initial read seems to be something related to layer reuse but I just skimmed things so far.

yes, that was surprising to me too. It would be a big deal if they switched to an encoder-decoder model like the original transformer. But I don't think that's what it's doing. One thing is the causal part, so in the original transformer, the encoder was bidirectional, but in this case it is not, so that's one difference. So I think it's an optimization for the prompt/prefill so that the attention is summarized into the output of the encoding layers, rather than all the layers. I just skimmed the paper too so if anyone else has insight, please correct me.

Also mid-range, A57 uses Xclipse 550 (RDNA 3) and A37 uses previous generation Exynos with Xclipse 530 IIRC.

Very interesting project, I like it. Just wanted to clarify though the sentiment analysis is just the count of stripped words and used to tag things with the emoji? I was initially expecting it to be a part of the actual command construction process (even though I couldn't figure out how that would be relevant) given how it was listed.

Ciao, yes for now the sentiment analysis is used only to provide an emoji related to the response. In the future I would like also to influence the choices of adjectives and interjections according to the sentiment.

For now it is a bit of a gimmick I agree :)


I feel like RLHF has a pretty obvious ground truth, human feedback is used as an (albeit noisy) signal of average human preferences. Same thing with RLVR and "solving the problem".

To be more specific there’s no ground truth tokens to predict. There a verifiable answer in RLVR. But the tokens are explored. Not predicted as there’s no true token to predict.

Diffusion LMs denoise a canvas which I personally find more interesting.

I don't really disagree that human cognition is essentially a predictive task though, as I understand it, predictive coding and related theories based on the Bayesian brain hypothesis are fairly popular these days (though maybe not clearly dominant over alterative models? IDK I'm not a neuroscientist). I imagine most people would draft a few tokens before refining them like MTP or diffusion though, if we do decide to use LMs as an analogy to human cognition.


Probably Fox would be the server with the most players but at DDK basically any would work and I think getting exposure to different playing styles would be worthwhile.

Maybe not to the extent you mentioned earlier (I definitely agree that it's not "focused" on this) but there is definitely asymmetric playout training in the training data (see comments on lightvector/KataGo issues #39 and #162 mentioning it for example), that is presumably how they got the tweak for playoutDoublingAdvantage (i.e. actually having a few thousand of the millions upon millions of training games be games where playouts have been doubled for advantage).

Yeah, thanks, I forgot that existed. That definitely weakens my point quite a bit.

It's still not _quite_ the same thing because an advantage in playouts is not a great model for how a stronger/weaker player dynamic actually works, but it is something for sure, way more than the nothing that I said exists.


AU user also, just checked AI studio since that seemed like the best bet and both 3.8 and 3.7 show up (and can be used for chat in playground, though IDK what the limits for that are). Chat in gemini.google.com is also 3.6 for me but I'm on free tier lol so I don't exactly expect it to show up any time soon. I think there's also another free API beyond the AI studio one (which is 20 RPD free according to docs so not really useful) but I forgot where it was (Google cloud maybe?) and what the limits for that were.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: