ImmanuelAI
Immanuel Peter

I like Artificial Intelligence.

Work
MTS Intern at Tensormesh
Study
CS, Physics, & Math at UChicago
Interests
Vision, Multimodal, World Models, Physical AI
Hugging Face
19,147 downloads · 18,118 on datasets, 1,029 on models

Work

Photo of a bald eagle used as the probe inputPCA of DINOv2 patch tokens: the eagle separates as one regionPCA of Muse Glimmer patch tokens: near-uniform speckle

One photo through two frozen towers. DINOv2 separates the bird as one region; Muse Glimmer shows no subject.

Vision Tower Bench

2026

A benchmark for evaluating vision head capabilities

AutoMoE architecture: four experts, a top-2 gate with context, and a trajectory head producing ten waypoints and speed

Four experts feed a context-conditioned gate that keeps the top two. The gate weights drawn are illustrative.

AutoMoE

2025

A MoE self-driving model in PyTorch

Fencing-first failover sequence in the Redis operator

Redis Operator

2026

CloudNativePG for Redis

Illustration of the Postplan dashboard

Postplan

2026

Easy toolchain to upload static files

Experience

Tensormesh

Mar 2026 – now

Member of Technical Staff Intern · Foster City, CA

  • Contributed upstream to LMCache, adding a coordinator endpoint for listing pinned cache entries and per-key access tracking in the key directory.
  • Shipped 35 PRs and 100+ commits across 9 company repos spanning LLM observability and inference optimization.
+ More− Less
  • Integrated Arize Phoenix into Tensormesh’s observability stack, emitting OpenInference LLM traces from the router for all requests, and shipped SDK and CLI tooling for inspecting traces and spans.
  • Reworked the Prometheus metrics for TTFT, throughput, and other rollups that power the serverless observability dashboards.
  • Built an orchestrator microservice that pins prompts in LMCache asynchronously during inference, and exposed pin management through the API, SDK, and CLI.
  • Built a KV cache discovery interface on top of the LMCache Coordinator API to browse cached keys, tokens, and utilization across engines.
  • Led an internal research effort on key weighting for long-context inference, from evals to writeups.

Quantum Rings

Jun – Aug 2025

Software Engineer Intern · Chicago, IL

+ More− Less
  • Delivered 19 PRs, 43 contributions, and 15 completed GitHub issues across the internship, adding ~15K LOC and removing ~3.6K LOC while reviewing code and driving schema refactors.
  • Migrated execution data from the user entity to a dedicated relational table with FKs, modularizing schema and ensuring test suite stability with no performance regression.
  • Implemented a telemetry aggregation background worker (AWS SQS + TypeORM) to asynchronously roll up user execution activity, improving scalability and simplifying downstream analytics queries.
  • Designed and deployed queue-driven execution processing to decouple heavy telemetry operations from the API, reducing request latency and enabling horizontal scaling.
  • Built full-stack admin analytics dashboards with NestJS, Next.js, and Recharts, integrating SQL time-bucket aggregation and timezone-safe filtering to track user growth, active usage, and execution volume.

Open source

LMCache

  • Added a coordinator endpoint for listing pinned cache entries. #4960
  • Added per-key access tracking to the coordinator key directory. #4873

vLLM Production Stack

  • Added router volume and mount support for read-only root filesystems. #975
  • Corrected the default NVIDIA runtime class, with regenerated CRD manifests. #974

Brev CLI

  • Authored the core rsync-first file-transfer implementation, with automatic SCP fallback and unit coverage. #423

Pyrefly

  • Upstream Rust cleanup standardizing the error-summary module and imports from “summarise” to “summarize”. #1370

Education

University of Chicago

2024 – 2028

B.S. Computer Science, B.A. Physics, B.A. Mathematics