Introduction
When doing recruiting or network-operator work, GitHub's public data is the cheapest source of "talent signal" — but it is not directly usable. A raw /users/<owner>/repos response has hundreds of fields and is nearly 250KB on a single page, and a browser hitting api.github.com directly also has to clear CORS. What lume-talentlens sets out to do is concrete: turn a person's GitHub footprint into decision material for "is this person worth following / worth hiring" (job-seeking status, engineering output, technical focus, network quality).
Its most counter-intuitive point is this: the server side is not Node/Python, but a single binary compiled from a self-designed DSL (Lume). This article does not walk you through building it from scratch; instead it tears the real code apart layer by layer — looking at how it uses Lume's routing / SSR / agent tools to cram "read GitHub snapshots, compute talent signals, run a chat agent" into one binary, while plugging the two pits of privacy and response-body size in advance.
TL;DR
- The server is a Lume single binary (prefork ×4), and
app/github.lumeonly does assembly:importeachlib/*.lumemodule to complete route/tool registration, thenserver {}+run()to start the service. - Two data sources, one serving path: by default it reads local offline snapshots (
data/github/<owner>/, zero network); only uncached owners go through the server-side/api/live/githubproxy to GitHub's API, bypassing the browser's CORS. - When the project started, Lume had no
lower/containsbuiltins yet, soshared.lumeshipped those two helpers (based onreplace); lume-core merged them into the language builtins on 2026-10-08, and talentlens keeps the local helpers only for backward compatibility.importdoes not propagate mutations across modules, soanalyze.lumeis all explicit(repos, snap)pure functions. - Response bodies have a hard cap (a historical release build would collapse at 64KB), so the whole project uses three unified tricks: aggregation (
/api/overviewreturns only stats), pagination (/api/repos≤50 per page), and projection (the live proxy slims the raw JSON down to consumed fields). - Safety gates rely on convention, not luck: bind only
127.0.0.1,.envfully gitignored, and Lume'senv()masks variables containingTOKEN/SECRET(the app reads the aliasGH_ANALYZER_PAT); write operations accept onlyapplication/jsonto block CSRF.
Overall Architecture: The Browser Trusts Only the Local Service
The topology is flat — the browser talks only to the local Lume service, and outbound traffic has only two destinations: the GitHub REST API and the LLM endpoint.
┌─────────────── Browser (127.0.0.1:8091)───────────────┐
│ React SPA (/, /chat) SSR pages (/overview etc)│
└───────────────┬───────────────────────┬───────────────┘
│ same-origin JSON / SSE │ same-origin HTML
┌───────────────▼───────────────────────▼───────────────┐
│ Lume server (single binary, prefork×4) │
│ entry github.lume → lib/* (routing/tools/rendering) │
│ bind 127.0.0.1:8091 │
└───┬───────────────┬───────────────────┬───────────────┘
│ read local file │ outbound http_get() │ LLM bridge (agent)
▼ ▼ ▼
data/github/<owner>/ api.github.com OpenAI-compatible endpoint
Key design: the server reads local files by default. Snapshots are prefetched offline by scripts/fetch-github.sh into data/github/<owner>/, and the long-running service does not actively go online. Only when a user queries an owner not yet cached does the server proxy through /api/live/github — this step simultaneously bypasses the browser's egress/CORS restrictions on api.github.com.
Server Side: Wiring Everything with a Lume Single Binary
The entry app/github.lume is deliberately thin; routes and tools live in lib/, and take effect through top-level registration. server_port() / server_bind() read the port and bind address from environment variables, defaulting to 8091 and 127.0.0.1:
import "lib/shared.lume" as sh;
import "lib/api.lume" as api;
import "lib/live.lume" as live;
import "lib/ssr.lume" as ssr;
import "lib/tools.lume" as tools;
import "lib/actions.lume" as actions;
server {
port = server_port();
workers = 4;
bind = server_bind();
docroot = "./www";
views = "github";
}
run();
Lume's module mechanism is the foundation that makes this split work: get / tool written at a module's top level complete route/tool registration on import by the entry point, with no explicit call needed; nested import is relative to the module's own directory (lib/shared.lume writes import "analyze.lume" rather than lib/...); export func is callable across modules. But the importer's mutations to an incoming map/array do not propagate back to the exporter — so the pure-computation layer must explicitly receive and return data, rather than mutating the passed-in reference.
Dependencies point in one direction: github.lume → each lib/*; analyze.lume depends on nothing; ui.lume only does SSR rendering.
Data Model and the Pure-Function Layer
The snapshot data/github/<owner>/snapshot.json is a pipeline artifact (the server reads it only). The raw API's ~100 fields are projected by fetch-github.sh down to ~24 fields/repo, and derived fields like created_year, recency buckets, and push_month are precomputed in the pipeline — the server only aggregates, never recomputes, which is why interfaces return in milliseconds.
lib/analyze.lume is the pure-function layer: every function takes (repos, snap) and returns JSON, with no IO, no env, no routing. Note that analyze builds a complete all array (with the full repo set), while what is exposed externally is analyze_lite — which drops all and keeps only aggregates and top-10, specifically to stay under the response-body cap:
export func analyze_lite(repos, snap) {
let a = analyze(repos, snap);
return {
owner: a.owner,
fetched_at: a.fetched_at,
profile: a.profile,
count: a.count,
non_fork_count: a.non_fork_count,
totals: a.totals,
languages: a.languages,
recency: a.recency,
years: a.years,
top_by_stars: a.top_by_stars,
top_by_activity: a.top_by_activity,
talent: a.talent,
push_trend: a.push_trend,
};
}
insight() computes talent signals from a recruiting perspective on top of analyze: the active_within_90d ratio, the with_desc/with_license/with_topics ratios, top2_language_share, and so on — these are exactly the same口径 shared by both the dashboard and the agent tools.
Owner Resolution and Snapshot Reading: The Security Boundary Starts at the Parameter
Every route/tool's first step is to resolve the owner. sanitize_owner in lib/shared.lume guards the owner as a single path segment: length ≤39, and forbidden characters . / % space # ? etc. The purpose is direct — the data/github/<owner>/ directory lookup can never escape data/github/:
export func sanitize_owner(s) {
if (s == null) { return ""; }
s = str(s);
if (s == "") { return ""; }
if (len(s) > 39) { return ""; }
if (contains(s, ".") or contains(s, "/") or contains(s, "%")
or contains(s, " ") or contains(s, "#") or contains(s, "?")
or contains(s, "\n")) {
return "";
}
return s;
}
export func load_snap(owner) {
let raw = read_file("data/github/" + owner + "/snapshot.json");
if (raw == null) { return null; }
let s = json(raw);
if (s == null) { return null; }
return s;
}
When the project started, Lume had no lower / contains builtins yet, so shared.lume shipped those two helpers, both cobbled together with replace (contains judges containment by "the string gets shorter after replacement"); lume-core merged lower/contains (along with upper/capitalize/trim/split/join/substr) into the language builtins on 2026-10-08, and talentlens keeps the local helpers only for backward compatibility. The multi-owner resolution priority is: ?owner= > OWNER env var > data/github/last_owner > default (when not found, returns empty string, so the JSON route goes to no_snapshot); when the snapshot is missing, the JSON route uniformly returns {status:404, error:"no_snapshot", owner:<owner>, hint:"run: OWNER=<owner> make fetch (or make fetch OWNER=<owner>)"}.
Live Proxy and the Response-Body Cap
/api/live/github in lib/live.lume is the server-side exit that proxies GitHub. It first passes live_safe_path, an SSRF guard — a fixed https://api.github.com prefix plus a path allowlist, so p can never smuggle in a different host or scheme:
func live_safe_path(p) {
if (p == "" or len(p) > 512) { return false; }
if (not sh.contains(p, "/users/")) { return false; }
if (sh.contains(p, ":") or sh.contains(p, "//") or sh.contains(p, " ") or sh.contains(p, "..")) { return false; }
return true;
}
A raw 100-repo page is ~250KB, far exceeding the framework's single-response-body cap, so slim_body must project before returning (/repos → 16 fields/item, /followers|/following → 5 fields, profile → 15 fields). The proxy also does two layers of caching: an in-memory TTL (LIVE_CACHE_TTL_SEC, default 300s) and a disk-backed data/github/live/ (shared across restarts, default 24h) — only successful responses are cached, so transient errors are never pinned. gh_get itself carries 3 retries (against the occasional TLS handshake reset on mainland networks), but 403/429 return immediately, no retry, no quota burn.
Frontend: A Lightweight React + esbuild SPA
The frontend is React 18 + TypeScript, bundled with esbuild, no webpack/vite, with pnpm managing dependencies. scripts/build-ui.sh bundles frontend/src/main.tsx into www/github/app.js. api.ts uniformly wraps fetch with a j<T>() layer: on 404/non-ok it tries to parse the no_snapshot body, and only throws if it cannot — because "no snapshot" is a normal business state and should not be treated as an exception:
async function j<T>(url: string): Promise<T> {
const r = await window.fetch(url);
if (r.status === 404 || !r.ok) {
try {
return (await r.json()) as T;
} catch {
throw new Error(url + " -> " + r.status);
}
}
return (await r.json()) as T;
}
The component tree is organized from an HR perspective: Dashboard.tsx assembles panels (overview/talent-signal/comparison), lists (followers/following/bots/worth-following + radar-profile cards), bars (language/activity/year bar charts); Agent.tsx takes the SSE route to /chat; LangSwitch + i18n.ts do the Chinese/English switch, with all copy going through a dictionary and the switch persisted. Lists are ranked by the server-side scores.json influence; high scores get a badge, and the bot chip's tooltip shows the judgment basis directly.
Agent Tools: 10 repo_*/github_*
lib/tools.lume registers 10 agent tools (repo_insights / repo_search / repo_language / repo_recency / repo_year / repo_stats / github_story / github_owners / github_radar / github_people). They run on forked workers, reading the snapshot fresh on each call, so there is no shared mutable state. github_people also merges the bot markers from suspects.json at the tool layer into the return, so the agent sees the same chips as the UI. /chat goes through Lume's LLM bridge; when LLM_* is not configured, a built-in offline fallback engine answers without a model dependency.
A Few Gates on Privacy and Security
This is not "add security after launch" — it is writing conventions into the architecture:
- Network exposure: bind only
127.0.0.1, no external listening; containers set0.0.0.0and publish only on the host loopback. - Credentials not in code:
.env,data/github/,logs/are all gitignored;scripts/env.shsuppressesbash -xwhen loading to prevent echo-back. - Credentials not in the app: Lume's
env()masks variables whose names containTOKEN/API_KEY/SECRET/PASSWORD; the live proxy reads the aliasGH_ANALYZER_PAT, which the Makefile maps fromGH_TOKEN. - Responses leak no secrets:
/api/github_authreturns only{authed:bool}; live/refresh responses contain no token. - Write-operation CSRF: follow/unfollow accepts only
Content-Type: application/json; cross-site JSON is blocked by the missing CORS header + preflight, and HTML form / text-plain payloads are 403-rejected before the body is even parsed. - Data isolation: snapshots/network/radar all live in the gitignored
data/github/; the code hardcodes no owner.
One pit worth recording: the public release binary of Lume (agent-httpd 1.0) does not build in http_get/put/delete, so running this app would fail at runtime — you must use the full build produced by make under work/lume/lume.
FAQ
Why does lume-talentlens's server use the self-designed Lume DSL instead of Node/Python?
Because it has to fit into a single binary: a native HTTP server + SSR + agent tools + LLM bridge in one stop, plus offline-only local snapshots and zero third-party Python dependencies. Lume’s top-level registration (effective on import) and server{} config make the split of routing/tools/rendering very lightweight. The cost is that when the project started Lume had no lower/contains builtins yet (merged into the language on 2026-10-08, with talentlens keeping the helpers for compatibility), and that the public release binary lacks outbound-HTTP builtins — these two constraints.
How exactly does the response-body cap affect the interface design?
The Lume framework has a hard cap on a single response body (a historical release build would collapse at 64KB). The whole project handles it with three tricks: /api/overview returns only aggregate stats, not the repo list; /api/repos paginates, ≤50 items per page; /api/live/github projects the raw GitHub JSON down to the app’s consumed fields (/repos → 16 fields/item) before returning. An unprojected 100-repo page at ~250KB would blow it up directly.
How does the browser bypass GitHub's CORS?
It does not hit api.github.com directly from the browser. Uncached owners are proxied by the server through /api/live/github; the server’s outbound path has none of the browser’s CORS/egress restrictions; live_safe_path fixes the host + path allowlist to prevent SSRF, and slim_body projects and slims before returning.
For a local-only recruiting dashboard, how is data privacy guaranteed?
Snapshots are pulled offline by a script into local data/github/<owner>/, and the long-running service reads locally only, zero network by default; .env and data/ are gitignored, so a public clone carries no personal data; Lume’s env() masks variables containing TOKEN/SECRET, and the app reads the alias GH_ANALYZER_PAT; /api/github_auth returns only a boolean, never a token; write operations accept JSON payloads only to block CSRF.
Do the 10 repo_*/github_* tools behind /chat read live data?
No, not live. They run on forked workers and read the local snapshot fresh on each call (data/github/<owner>/snapshot.json), so there is no shared mutable state and they work offline. To refresh data you must first OWNER=x make fetch or hit /api/refresh (rate-limited to 60s per owner), and only then will queries see the new snapshot.
make dev won't start and reports an http-related undefined — what do I do?
That means you are using the public release binary (agent-httpd 1.0), which does not build in http_get/put/delete. Switch to the full Lume build produced by make under work/lume/lume (specify the path with LUME=). The make dev startup probe validates the outbound-HTTP builtin with printf 'print(http_get);' | $LUME -, and refuses to start if it is missing.
Project Links
- GitHub repository: https://github.com/erishen/lume-talentlens
| Module | File | Notes |
|---|---|---|
| Entry & assembly | app/github.lume | import modules + server{} + run() |
| Shared layer | app/lib/shared.lume | owner resolution, snapshot reading, string helpers, gh_get retries |
| Pure analysis layer | app/lib/analyze.lume | repo_view / analyze / insight / histograms (no IO) |
| JSON routes | app/lib/api.lume | owners/overview/repos/people/radar/refresh and 16+ endpoints |
| Live proxy | app/lib/live.lume | /api/live/github proxy + SSRF guard + two-layer cache |
| Agent tools | app/lib/tools.lume | 10 repo_*/github_* agent tools |
| Frontend API client | frontend/src/api.ts | unified fetch wrapper + no_snapshot handling |
| Architecture doc | ARCHITECTURE.md | layering, data model, data flow, security boundaries |
Leave a reply