spacestr

🔔 This profile hasn't been claimed yet. If this is your Nostr profile, you can claim it.

Edit
someone
Member since: 2022-12-22
someone
someone 1d

maybe one of the hard parts of my kind of fine tuning is evals. where do you want model to go? how do you find true answer of a hardly debated issue? instead of manually writing answers in many domains (which i can't do, i don't know many answers in many domains) i rely on a few tricks: - find other aligned llms and get ideas from them - rank many llms in AHA leaderboard and get ideas from top ones also rejecting the worst ones - do mixture of agents of the above to get a collective answer - and lately, compare the answers coming from my fine tunes with base models, assume my fine tune is preferred if there is a difference these still can't find perfect answers but they kick the model in the right direction. and that may be a big deal. can't claim my fine tune (ostrich) knows every truth. but it may make more sense to claim if the base differs with fine tune most probably fine tune is the better answer. making better evals ends up training better models. and produce better AHA ranking. which further sharpens evals. this feedback loop is going to be useful for a while.

someone
someone 6d

i've been fine tuning llms for 2 years now. mostly did qlora. recently experimenting with 'behavior steering' where instead of spending hours making a lora adapter, you try "brain surgery". these are like quick math operations to change behavior of an llm. you can install things like bitcoin lover, herbalist, fasting lover. turns out all of these personas have different difficulty levels. you can easily install fasting lover because qwen 3.5 and 3.6 doesnt resist it (in other layers). an interesting finding today, guess which persona is hardest or sometimes impossible to install: vaccine hater!

someone
someone 6d

getting ready to fine tune 3.8 - evolutionary strategies - behavior steering experiments - expanded dataset - bringing back ORPO - more orthogonal evals to keep overfitting minimum - most probably will take abliterations as base, either mine or somebody else's - random entropy addition from huggingface fine tunes (take what is popular on hf and randomly introduce into the lineage) - bring more vibe coding: turns out LLMs know how to fine tune

someone
someone 12d

the ostrich model can do contemplations well now. that means it can expand the tiny dataset that we have. this could theoretically end up making less overfit models with the same degree of AHA score. imagine an automatic alignment agent, that scans whatever is out there and consumes (trains) if classifies as true.. could be practical one day. we could let it run and auto train and self evolve. a truth db might still be needed, augmenting the reasoning and decisions of this agent.. that is harder to construct but we could.. this is good news. recursive self alignment might be here soon.

someone
someone 15d

woah they really started with HF! they must be loving their children so much 😉 https://www.wired.com/story/hugging-face-has-a-nonconsensual-deepfakes-problem/ https://www.engadget.com/2224899/hugging-face-reportedly-plagued-with-deepfakes/ https://www.theverge.com/ai-artificial-intelligence/971723/hugging-face-nudify-deepfake-undress-women-children time to decentralize LLM sharing...

someone
someone 1d

why are phone numbers and email addresses are kept in a shipping provider?

someone
someone 15d

there is a new version of ostrich, more successful with long context jobs. https://huggingface.co/etemiz/Ostrich-27B-260721 apparently when i push for higher AHA score it ends up overfitting and merging it with previous 3.5 version healed the overfittings. this gave me another idea: what if we merge all the fine tuners like fine tunes going for abliteration (uncensored) and also separately agentic coding. since the merging heals, everybody's overfittings can cancel each other. if that works, this could mean models that we create are like our overfitted ideologies. we humans all are like knowing and believing some stuff religiously and when or if we can get together those extremes kind of relaxed. socialization of humans look exactly like merging of LLMs..

someone
someone 16d

you can use torrents

someone
someone 18d

https://cdn.hzrd149.com/62c27d50f2843e369b0ebdb4d40025da99d0d9b54b90f80f483563044d583bd3.torrent

someone
someone 18d

its a good client. my 9 year old phone hardly runs it but my 2 year old likes it a lot. experience is faster than primal. primal was supposed to be fast with the cache etc but the server is in germany and round trip to usa looks slow. and maybe they dont delete events. so end up responding slower than amethyst. amethyst maybe fetching from closer usa servers first and then deduping with german servers?

someone
someone 18d

aws banned me because I automated download of K3 on aws using GLM 5.2. my plan was to convert K3 to torrents within minutes.. is using AI to create instances and start scripts a bad practice?

Welcome to someone spacestr profile!

About Me

Interests

  • No interests listed.

Videos

Music

My store is coming soon!

Friends