978 Downloads? 🤔

#12
by cimilarkes - opened

Vram rich people share your specs :D

I don't think its real people, I think its bots.

Actually - i really shouldnt say this... but streaming downloads, ie streaming huge datasets... count each shard streamed as one download...
ie - 50,000 shard db = 50000 downloads per full streaming. What i think it is is some provider company streaming the model to their servers to run it.

thats just from my personal experience tho.

Yeahhhh, Bots...
image

why would a bot download like... 6tb repo?

likely a bunch of package manager solutions as well as clones of HF. Also likely a bunch of benchmark sites and model providers.

i run it on my samsung s24 FE locally with termux idk the specs tho but it work fine 200 toks

i run it on my samsung s24 FE locally with termux idk the specs tho but it work fine 200 toks

Proof

i run it on my samsung s24 FE locally with termux idk the specs tho but it work fine 200 toks

Proof

how to take picture of phone while phone is currently being used

so actually - its possible... technically. To do it you have to use something like colibri (github project to run glm on 24gb) but like... stream each shard when needed and then delete it.

i looked up the specs - 8gb in ram and 128gb disk - so its technially possible... but very very slow.

so actually - its possible... technically. To do it you have to use something like colibri (github project to run glm on 24gb) but like... stream each shard when needed and then delete it.

i looked up the specs - 8gb in ram and 128gb disk - so its technially possible... but very very slow.

What would the approximate tokens per second for this be on a Nokia?? I'm thinking it's close to one token every 6 months maybe?

so actually - its possible... technically. To do it you have to use something like colibri (github project to run glm on 24gb) but like... stream each shard when needed and then delete it.

i looked up the specs - 8gb in ram and 128gb disk - so its technially possible... but very very slow.

its super fast for me idk bro

so actually - its possible... technically. To do it you have to use something like colibri (github project to run glm on 24gb) but like... stream each shard when needed and then delete it.

i looked up the specs - 8gb in ram and 128gb disk - so its technially possible... but very very slow.

What would the approximate tokens per second for this be on a Nokia?? I'm thinking it's close to one token every 6 months maybe?

it get 100 toks for me not that fast like datacenter fast but it fast for locally

well then you prob aint running it on a phone lol... maybe local llm setup? since its a larger model id say smth like REAP (remove un needed experts) + dgx station?

well then you prob aint running it on a phone lol... maybe local llm setup? since its a larger model id say smth like REAP (remove un needed experts) + dgx station?

i would run it on my laptop but i broke my school laptop so i only have samsung phone right now sometimes when termux is slow i just put the model into base64 and put it inside of a html becauae html has easier access to my phones gpu and run it faster 320 avg toks but i use python in termux for agenticness

Sign up or log in to comment