Menu

#Token

415 posts

Feed·
20 of 415 posts
GitHub - argonautlabsai/deltafin: ARGODRIVE Deltafin: Kimi K3 (2.8T MoE) streamed from SSDs on Apple Silicon — fork of gavamedia/deltafin with the ARGODRIVE storage work and benchmark package
🖼️
0

GitHub - argonautlabsai/deltafin: ARGODRIVE Deltafin: Kimi K3 (2.8T MoE) streamed from SSDs on Apple Silicon — fork of gavamedia/deltafin with the ARGODRIVE storage work and benchmark package

Hacker News·Hacker News·7 days ago
#GbimiaTR
#github#deltafin#token#full#release#prompt

ARGODRIVE Deltafin: Kimi K3 (2.8T MoE) streamed from SSDs on Apple Silicon — fork of gavamedia/deltafin with the ARGODRIVE storage work and benchmark package - argonautlabsai/deltafin

15s
Read More
Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems
🖼️
79

Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems

#arxiv#context#cost#token#view#production

Production AI agents' failures are less often due to an inability to reason well and more often because they cannot manage what is in their reasoning context: conversation histories, large prompts, large tool definitions, and ballooning tool outputs.…

15s
Read More
GitHub - sqliteai/waste: Run the full 2.78-trillion-parameter Kimi K3 model beyond available RAM by streaming activated weights directly from NVMe. A dependency-free, embeddable C inference engine.
📰
0

GitHub - sqliteai/waste: Run the full 2.78-trillion-parameter Kimi K3 model beyond available RAM by streaming activated weights directly from NVMe. A dependency-free, embeddable C inference engine.

Hacker News·Hacker News·about 2 months ago
#i9E3YRcs
#github#container#waste#cache#token#engine

Run the full 2.78-trillion-parameter Kimi K3 model beyond available RAM by streaming activated weights directly from NVMe. A dependency-free, embeddable C inference engine. - sqliteai/waste

15s
Read More
 The enduring paradox of the AI economy — models get better and more efficient, yet costs can still easily spiral out of control
📰
0

The enduring paradox of the AI economy — models get better and more efficient, yet costs can still easily spiral out of control

Latest from Tom's Hardware ·Bruno Ferreira·about 2 months ago
#nUYzH3aI

Token amplification creates a paradox in the AI economy, as more capable models beget more complicated tasks.

15s
Read More
Qwen3.8 is launching and going open-weight soon!🌐

With a massive 2.4T parameters, this model is continuously evolving. We believe it’s one of the most powerful model available today, compatible to leading frontier AI models , second only to Fable 5.

You don't have to wait to https://t.co/JS3ID73IYS
🖼️
0

Qwen3.8 is launching and going open-weight soon!🌐 With a massive 2.4T parameters, this model is continuously evolving. We believe it’s one of the most powerful model available today, compatible to leading frontier AI models , second only to Fable 5. You don't have to wait to https://t.co/JS3ID73IYS

Hacker News·Hacker News·about 2 months ago
#6hPsiVYT
#twitter#token#qwen3#model#wait#plan

Qwen3.8 is launching and going open-weight soon!🌐 With a massive 2.4T parameters, this model is continuously evolving. We believe it’s one of the most powerful model available today, compatible to leading frontier AI models , second only to Fable 5.…

15s
Read More
The AI tokenmaxxing party is crashing over spiraling costs — leaked consulting firm audio suggests no one is sure…
🖼️
0

The AI tokenmaxxing party is crashing over spiraling costs — leaked consulting firm audio suggests no one is sure…

Latest from Tom's Hardware ·Jon Martindale·3 months ago
#XwnSSISf

"Leadership [...] are still asking the question of whether they're getting value from what we're spending."

15s
Read More
 AI costs spike as subscriptions hit pricing wall — firms turn towards Chinese LLMs, open-source models to extend budget
📰
0

AI costs spike as subscriptions hit pricing wall — firms turn towards Chinese LLMs, open-source models to extend budget

Latest from Tom's Hardware ·Jowi Morales·3 months ago
#g0lcpl7G

Companies look for cheaper alternatives as token costs for frontier AI models skyrocket, potentially impacting OpenAI and Anthropic's bottom lines.…

15s
Read More
Token-Savior's 5 Hidden Uses: The MCP Server That Cuts Your AI Coding Costs by 80%
🖼️
0

Token-Savior's 5 Hidden Uses: The MCP Server That Cuts Your AI Coding Costs by 80%

DEV Community··3 months ago
#lXPpUo1r
#dev#token#claude#savior#code#hidden

You probably use Claude Code, Cursor, or Windsurf every day — but if you're not running an MCP server...

15s
Read More