P-Product

Software Engineer โ€” RISC-V SIMD Kernels for LLMs

$$$$

About the Project

The team builds low-level compute kernels for inference/training of well-known LLM models, hand-optimized for a proprietary (custom) chip based on the RISC-V architecture with RISC-V Vector Extension (SIMD intrinsic extension) and a custom VCIX Coprocessors. In terms of functional scope, this is analogous to what PyTorch/Tensorflow provides (matmul/GEMM, convolution, elementwise ops, reduction, attention/softmax, quantization) โ€” but written from scratch for a specific chip rather than routed through an existing backend.

 

What You Will Actually Do

  • Write new compute kernels (functionally equivalent to PyTorch operators: matmul/GEMM, convolution, elementwise, reduction, attention/softmax) targeting custom RISC-V hardware.
  • Implement and tune mixed-precision inference pipelines across multiple numerical formats (FP32, BF16, MXFP8, MXFP4, NVFP4).
  • Investigate and improve the performance of existing kernels and how they are invoked: optimize cache vs. DRAM memory usage (memory hierarchy, tiling/blocking, prefetching, bandwidth vs. latency trade-offs).
  • Optimize distributed communication for multi-device and multi-node inference.
  • Partition and schedule workloads across compute cores to maximize hardware utilization.
  • Fuse individual kernels into efficient end-to-end pipelines for LLM model execution.
  • Write C++23 code with heavy use of metaprogramming (templates, constexpr, concepts, CRTP, SFINAE) to generate specialized code for different kernel configurations.

Required Qualifications

  • Hands-on experience writing low-level compute kernels โ€” GEMM/matmul, convolution, elementwise, reduction, or similar โ€” regardless of whether this was for PyTorch/ATen, XLA, CUDA, ARM NEON/SVE, x86 AVX, or other hardware. The key requirement is actual kernel-writing experience, not just usage of high-level ML frameworks.
  • Experience with SIMD/SIMT/vector intrinsics on any architecture (RVV, ARM NEON/SVE, x86 SSE/AVX/AVX-512) or with inline assembly, understanding how the compiler maps intrinsics to instructions etc.
  • Solid understanding of memory hierarchy: cache vs. DRAM, cache lines, latency/bandwidth trade-offs, data locality, tiling/blocking for compute optimization.
  • Strong interest in and willingness to optimize code for performance and memory efficiency, with the ability to identify bottlenecks and continuously improve execution speed, memory usage, and overall system efficiency.
  • Strong focus on latency minimization in production inference paths โ€” able to identify and eliminate latency-critical bottlenecks across the full execution pipeline, not just individual kernels.
  • Comfortable working in a Linux-like environment for building/testing (WSL/Linux), and experience debugging without full GUI debugger tooling (log-based debugging).

 

Strong Pluses

  • Direct RISC-V experience (ISA, RVV / custom intrinsic extensions, cross-compilation with a RISC-V toolchain โ€” GCC/LLVM/Clang).
  • Experience optimizing kernels for LLM inference/training (attention, KV-cache, softmax, layer norm, quantized matrix multiplication).
  • Experience with proprietary/custom hardware (not reference architectures), including writing code for early-stage chips where tooling and debugging capabilities are limited.
  • Understanding of collective communication operations, distributed training (Data/Tensor/Pipeline Parallelism).
  • Understanding of floating-point formats โ€” FP32, BF16, MXFP8, NVFP4/MXFP4 โ€” including their trade-offs in precision, dynamic range, and performance.
  • Understanding of how numerical format choice affects throughput, memory bandwidth, and training/inference numerical stability.
  • Strong C++ (17/20/23) with confident use of template programming โ€” real experience in a generic/template-heavy codebase, not just familiarity with the syntax.

Required languages

English B2 - Upper Intermediate
Ukrainian B2 - Upper Intermediate
Published 20 July
10 views
ยท
1 application
Response activity: Very high
Last responded yesterday
To apply for this and other jobs on Djinni login or signup.
Loading...