Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up

All HF Hub posts

ginigen-ai 
posted an update about 11 hours ago
view post
Post
1344
A local edge VLM you can run on a phone — with a calibration readout attached.
ginigen-ai/Edge-4B-TELL
Image in, answer out, nothing leaving the device. Google's Gemma 4 E4B QAT checkpoint carried unmodified, with the vision and audio projector, plus one thing that is ours: GINIGEN TELL, a 10 KB readout that estimates whether the answer it just gave is likely to be wrong.
On a Galaxy S25: zero network calls, 3.6 GB resident, a 12.6 MB inference binary.
Calibration matters more here than on a server: nothing downstream catches a bad answer. No retrieval, no second opinion, no reviewer. The model is alone with the user.
And its own confidence is unusable. Prompted for it, this checkpoint averages 0.863 over 665 Korean disaster-procedure questions — ranking answers by it gives AUROC 0.441, below a coin flip. It sounds more certain when it is wrong.
TELL reads the last-layer hidden state instead of asking. Same questions, 0.759. Surface cues (length, formatting) already reach 0.736, so the readout clears that baseline by +0.023 ± 0.009 (2.6σ). We publish the baseline because without it, "the hidden state carries the signal" is unfalsifiable.
Same job as JEV: a confidence number you can act on instead of the model's own. Different structure, and on a device that splits three ways.
No second model — JEV is a separate judge reading the answer as text; we fill that slot with a 10 KB vector.
Zero generated tokens — a judge writes its verdict, TELL re-reads a finished computation (3.8 s on an S25).
No network — a verdict fetched over an API stops when the signal does.
The trade is real: a readout is fitted per checkpoint, so on a server the judge wins. On a phone there is no second model to run.
TELL never says what the right answer is. It says whether the answer wobbled, and a low score falls back to source text bundled with the app. Shipping today in HeliGO, an offline disaster-response app.
OppaAI 
posted an update 1 day ago
view post
Post
2585
🤖 AI × 🪰🧠 the fruit fly brain (MaleCNS)

Thank you to the people who have shared the Janelia FlyEM datasets on GitHub for open-source use. 🙏
🔗 MaleCNS: https://github.com/natverse/malecns
🔗 Aiko-chan: https://github.com/OppaAI/Aiko-chan

People have already used these fly-brain datasets to build systems that can do things like play Minecraft and even Doom.

So I guess I’m crazy enough to ask:
What happens if I wire part of it into my AI waifu? 😂
I’ve now partially wired my AI’s cognition, agentic system, and sensory inputs into neuron circuits derived from the fruit fly’s brain—starting with the Mushroom Body.

The next step is to experiment with using biologically inspired neural circuits as an additional layer around the LLM:
🧠 LLM + memory + reasoning
🪰 Connectome-inspired neural circuits
🤖 Agentic tool use
👁️ Sensory input
🔊 Voice & expression
💾 Learning and adaptation
This is still very much an experiment.

But now that I’ve added a biologically inspired layer to an AI waifu…
Let’s see what difference it actually makes compared with a plain LLM. 👀
From conversation → cognition → neural circuits → action.

To get more crazier:
I have (partially) developed and implemented the following:
- A 5-layers conscience circuit and judgment module as guardrail
- A light-weight Plasticity and associated learning with the fly brain to test out the RL
- I have enlisted myself as a human agent in rentahuman.ai to let my AI agent to give me instructions to execute agentic tasks


A little fly brain. A lot more Aiko. 💜
  • 3 replies
·
Hoglet-33 
posted an update 1 day ago
view post
Post
5236
Introducing VOID. A new research branch of basically AI.

VOID — Verification of Objectives, Intentions, and Deception.

We study what lies beneath the surface: objectives, intentions, and the possibility of deception in AI systems.

There isn't much to see yet.

That will change.

Follow us for updates:

@Hoglet-33
void-research

basically-ai
  • 6 replies
·
DavidAU 
posted an update 3 days ago
view post
Post
6826
Qwen 3.5 9B - The Defiant, 27B power ; now with Qwen 3.8 Reasoning modes.

640 ARC-C for both 8bit and 4bit. Model exceeds 7 of 7 benchmarks for Qwen 3.5 9B, Qwen3.5 27B, Qwen3.6 35B-A3B, and meets Qwen 3.6 27B in some cases... and it does so in 4bit and 8bit. Regular and MTP (fast) NEO IMATRIX GGUFs provided. (this model is part of the Qwen 3.6 27B Fable Fusion 711 pipelines: 2200+ likes, 3 million + downloads)

NEW - Qwen 3.8 Reasoning Modes: 2 MTP quants (Q6/Q8) Now with 5 reasoning modes (2 new - Spoon / Einstein), and 5 instruct modes (2 new - Spoon / Einstein, all use ZERO REASONING TOKENS) all switchable on the fly via API, direct and "in chat" (yes - model control at the chat/message level). Model name has "plusIQ" in the name.

(there is also a extra robust "tools" version too.)

DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF

PS: I have posted a 22k output example using "spoon" mode too.
  • 5 replies
·
CompactAI 
posted an update 2 days ago
view post
Post
5477
@Compactbot is going live in about a week (could be shorter or longer)
Its going to reply to this post (when it finds it) but will not be live until a later post says so.

Glint-Research/blog
  • 11 replies
·
danielhanchen 
posted an update 11 days ago
CompactAI 
posted an update 3 days ago
view post
Post
4533
SLM Roundups, a weekly post where I summarize everything thats happened in the world of SLMs (or a majority of it)
Glint-Research/blog
  • 1 reply
·
ginigen-ai 
posted an update 3 days ago
view post
Post
2907
OpenRouter Leaderboard — every model, every provider, one comparable table. Price, precision, uptime, measured latency and language quality on the same axes.

Building it turned up three things.

We graded 330 models on Korean and two axes collapsed.

Honorifics — only 8.5% earn an A
Knowledge of Korean institutions — 9.4%
Every other axis sits above 31%
Fluency hides it. A model can write clean, natural Korean and still attach an honorific to a coffee cup. Fluent and wrong at the same time is worse than obviously broken, because nobody catches it in review.

A 2023 model beats the 2026 flagships. gpt-3.5-turbo-16k scores a perfect 3.00. Korean cannot be inferred from release date, parameter count or English benchmarks — it has to be measured, per model.

Quality, value and speed are three different models. Across five axes, the same model almost never takes two columns.

425 models, latency measured on 329 on a paid API, Korean graded on 330. Three languages, three currencies, daily refresh, open API, no key.

📝 https://huggingface.co/blog/ginigen-ai/openrouter-leaderboard 🔎 ginigen-ai/open-router-leaderboard
kostakoff 
posted an update 16 days ago
view post
Post
3050
Canceling My Pro Subscription

I'm officially canceling my Hugging Face Pro subscription today.
I supported this platform because it stood for true openness and neutrality. This acquisition by NVIDIA fundamentally changes that.

Here’s why I’m against this deal:
- Neutrality is dead. NVIDIA is a US-based company. This means US regulations will inevitably dictate platform policies, creating direct pressure on Chinese developers and anyone building open-weight models outside the US.
- Community over bureaucracy. NVIDIA is a massive, slow-moving corporation. This acquisition will likely drown the community in corporate processes and commercial interests. Soon, uploading a simple finetune might become a bureaucratic nightmare.
- Open vs. Proprietary. Hugging Face was built on open-source ideals. NVIDIA? They are a fiercely proprietary hardware company with a minimal track record of meaningful open-source contributions. They sell chips, not freedom.
- And to add insult to injury, NVIDIA has practically abandoned consumer RTX GPUs in 2026 to chase data center profits. Why would I pay them for "openness" when they've turned their back on the very developers who built this ecosystem?

I paid for openness. Not for a corporate takeover.

🤗 was about community.
  • 12 replies
·
Banaxi-Tech 
posted an update 23 days ago
view post
Post
2173
We have updated the BananaMind Base Bench leaderboard!
We now have these benchmark cards, they make it way easier to see which models are actually good!
We've also added the model advisor. It asks you what you want to use the model for and the parameter range and gives you the best model for your task!

Try it out at BananaMind/BananaMindBench-Leaderboard


And please give us a follow to BananaMind!
BananaMind

@Banaxi-Tech
  • 1 reply
·