Menu

#Cuda

46 posts

Feed·
20 of 46 posts
GitHub - PJHkorea/photonic-mesh-fng-router: A humble exploratory PoC for a hardware-native, optical timing-frozen control plane engine. Fuses inline CUDA PTX, RAII memory tunnels, and 4D JAX shard_map structures to investigate 0ns-overhead fault-tolerant routing for hyperscale distributed AI.
📰
0

GitHub - PJHkorea/photonic-mesh-fng-router: A humble exploratory PoC for a hardware-native, optical timing-frozen control plane engine. Fuses inline CUDA PTX, RAII memory tunnels, and 4D JAX shard_map structures to investigate 0ns-overhead fault-tolerant routing for hyperscale distributed AI.

Hacker News·Hacker News·about 1 month ago
#sZXcuiUa

A humble exploratory PoC for a hardware-native, optical timing-frozen control plane engine. Fuses inline CUDA PTX, RAII memory tunnels, and 4D JAX shard_map structures to investigate 0ns-overhead f...

15s
Read More
GitHub - NVlabs/cutile-rs: cuTile Rust provides a safe, tile-based kernel programming DSL for the Rust programming language. It features a safe host-side API for passing tensors to asynchronously executed kernel functions.
🖼️
106

GitHub - NVlabs/cutile-rs: cuTile Rust provides a safe, tile-based kernel programming DSL for the Rust programming language. It features a safe host-side API for passing tensors to asynchronously executed kernel functions.

Hacker News·Hacker News·3 months ago
#tTZipU3j
#github#cutile#cuda#rust#tile#kernel

cuTile Rust provides a safe, tile-based kernel programming DSL for the Rust programming language. It features a safe host-side API for passing tensors to asynchronously executed kernel functions. -...

15s
Read More
 Nvidia's long-awaited N1/N1X SoC specs leak ahead of Computex launch — N1 to feature up to 20 Arm-based cores, standard N1 equipped with 12- and 10-core configs
📰
0

Nvidia's long-awaited N1/N1X SoC specs leak ahead of Computex launch — N1 to feature up to 20 Arm-based cores, standard N1 equipped with 12- and 10-core configs

Latest from Tom's Hardware ·Hassam Nasir·4 months ago
#xpN10TnF

The N1X reportedly comes in two SKUs: a top-end 20-core option with 6,144 CUDA cores matching the desktop RTX 5070, and a cut-down 18-core option with 5,120 CUDA cores.…

15s
Read More
NVIDIA CUDA 13.3 Enhances GPU Development with Tile Programming in C++, Compiler Autotuning, and Python
🖼️
0

NVIDIA CUDA 13.3 Enhances GPU Development with Tile Programming in C++, Compiler Autotuning, and Python

NVIDIA Technical Blog·Jonathan Bentz·4 months ago
#MjyCwVQT
#developer#include#cuda#cccl#import#python

NVIDIA CUDA 13.3 brings new capabilities and performance optimizations to developers across the CUDA ecosystem. The launch of NVIDIA CUDA Tile programming in…

15s
Read More
Develop High-Performance GPU Kernels in C++ with NVIDIA CUDA Tile
🖼️
0

Develop High-Performance GPU Kernels in C++ with NVIDIA CUDA Tile

NVIDIA Technical Blog·Jonathan Bentz·4 months ago
#VjvQQoBg
#developer#include#define#tile#auto#float

Developers can now use NVIDIA CUDA Tile programming within large existing C++ GPU codebases to develop highly optimized GPU kernels using tile-based…

15s
Read More
Why CUDA kernels silently corrupt memory and how to catch the bug
🖼️
0

Why CUDA kernels silently corrupt memory and how to catch the bug

DEV Community·Alan West·4 months ago
#pAAsZBHU
#ifndef#cuda#rust#kernel#scratch#compute

A practical guide to debugging silent memory corruption in CUDA kernels, with compute-sanitizer workflows and a look at Rust-on-GPU tooling.

15s
Read More
How to Deploy Llama 3.2 90B with vLLM + Speculative Decoding on a $16/Month DigitalOcean GPU Droplet: 2.5x Faster Inference at 1/110th Claude Cost
🖼️
0

How to Deploy Llama 3.2 90B with vLLM + Speculative Decoding on a $16/Month DigitalOcean GPU Droplet: 2.5x Faster Inference at 1/110th Claude Cost

DEV Community·RamosAI·4 months ago
#pACPdpCa

From Dev.to - tutorial: How to Deploy Llama 3.2 90B with vLLM + Speculative Decoding on a $16/Month DigitalOcean GPU Droplet: 2.5x Faster Inference at 1/110th Claude Cost

15s
Read More
CUDA Proves Nvidia Is a Software Company - Slashdot
🖼️
0

CUDA Proves Nvidia Is a Software Company - Slashdot

hardware.slashdot.org·hardware.slashdot.org·4 months ago
#ZvzQBICT
#comments#modal_box#cuda#nvidia#gpus#single

Nvidia's real AI moat isn't "a piece of hardware," writes Wired's Sheon Han. It's CUDA: a mature, deeply optimized software ecosystem that keeps machine-learning workloads tied to Nvidia GPUs.…

15s
Read More
The cuda-oxide Book — cuda-oxide
🖼️
0

The cuda-oxide Book — cuda-oxide

nvlabs.github.io·nvlabs.github.io·4 months ago
#4PcEYlZi

cuda-oxide is an experimental Rust-to-CUDA compiler that lets you write (SIMT) GPU kernels in safe(ish), idiomatic Rust. It compiles standard Rust code directly to PTX — no DSLs, no foreign language bindings, just Rust.

15s
Read More