Software Engineer – LLM Inference Optimization
Preferred Networks
Description
Preferred Networks (株式会社Preferred Networks) — Japan's leading AI unicorn, builder of the PLaMo LLM family and MN-Core AI processors, with research teams that have collaborated with Toyota, FANUC, and NTT on industrial AI — is hiring a Software Engineer to improve the inference engine powering its PLaMo API service and maintain PLaMo implementations in open-source projects including vLLM, based onsite in Tokyo, Japan. Preferred Networks will provide full visa sponsorship and relocation support for international candidates. This role focuses on optimizing how PLaMo models run in production — latency, throughput, cost — and contributing improvements upstream to open-source inference frameworks. It offers a rare opportunity to work at the frontier of Japanese-language AI while contributing meaningfully to the broader LLM inference ecosystem. The role was posted directly by Preferred Networks in the HN "Who is Hiring? (August 2026)" thread.
Required Skills
Similar Jobs
LLM Serving Engine Engineer
NEWPreferred Networks
2h ago
Salary not disclosed
LLM Serving Engine Engineer
NEWPreferred Networks2h ago
Salary not disclosed
Researcher
GiveWell
27d ago
USD 200K - 220K/yr
Researcher
GiveWell27d ago
USD 200K - 220K/yr
Senior Fullstack Engineer, Monetization
MoonPay
3mo ago
Salary not disclosed
Senior Fullstack Engineer, Monetization
MoonPay3mo ago
Salary not disclosed
Tech Lead Manager
Wheely
3mo ago
Salary not disclosed
Tech Lead Manager
Wheely3mo ago
Salary not disclosed
.png)