Hacker Newsnew | past | comments | ask | show | jobs | submit | ViscountPenguin's commentslogin

I doubt that any algorithms would be able to find a global maximum in such a large multidimensional non-convex space

Exactly correct, they are only capable of tasks that any school child could do; like solving millenium prize problems, hacking into tech companies, or tuning particle colliders. Nothing to see here.

There are humans behind all of these actions. We're just not holding them responsible for some reason re: hacking into tech companies.

"Unplug all possible target datacentres" has a pretty big blast radius, and people are unlikely to be coordinated or cooperative enough to actually do it.

The point of the paper isn't that embeddings contain information, it's that even if you don't know what model generated a set of embeddings you can still recover information from the geometry of the point cloud itself.

The fact that this is possible also adds some pretty strong restriction to the set of possible maps you could use to remove that information. No linear map will work since all embedding spaces are ~an orthonormal matrix apart, so some form of encryption is necessary. This wasn't known until very recently.


Ah right, thank you for clarifying.

But I think my point still stands, isn't the geometry information THE information I referred to in the first place? Obviously the vector size gives you the granularity but it's kind of unavoidable to positionally encode information in a latent space...that's literally what they're for?

But yes, it is very cool to know that regardless of exact implementation finding x,y,z representations of some dataset with various relationships (like language) creates similar geometry/clues across all the implementations.


At least for me, the Flatpack produces no sound.

The appimage also produces no sound :(

Unfortunately this approach doesn't feel that great down here in Australia, definitely a function of latency.

I think you could get a lot closer by framing this as an optimization problem, where you use the full alphabet dictionary, but add a residual prediction which aims to cover as much of the remaining domain name tree as possible weighted by popularity. This tree could then be pre-baked and stored with the same system. This would probably get you p99 0ms even in Australia.


Interesting, I didn't think of that yet.

I think most people search for domains they own, which doesn't correlate with Tranco popularity ranking. Treating 'popularity' not as a function of visitors, but as a function of number of known domain names with that prefix could work, though.


Fair enough, if people visit frequently you could probably save their previous searches as a cookie too.

A quick look at the continuous diffusion models linked in the post shows lots of transformer models still

This comes from the binomial expansion of (X+y)^n, where X^k y^(n-k) has coefficient (n choose k). This is since you "choose" X in K of the (X+y)'s (in the Bezier case it's kind of writing (t+(1-t))^n, the abcd are to not make it equal to 1).

The choice function is also exactly the same as the pascal triangle!


One of the really awkward points of stats is that not only does sample size matter, but also model specification. Very small misspecifications can easily lead to infinitesimal P values over large sample sizes.

Similar things are true of Bayesian stats, leading to things like predictively oriented posteriors being studied nowadays.


In my experience, a not insignificant chunk of students actively in a stats class think that a p value is the probability of the null hypothesis. My money is on more like 3% of educated professionals.


If you asked me for the definition of p value I could give it to you but if you asked me what it meant before this thread I would have said the same thing as you’ve said here.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: