Source: handbook/ in the repository. Every code block compiles: CI runs ray check on each one.
Performance
How to make a raylang program fast, in the order it pays to do it: measure, find where the time goes, fix the algorithm, compile to native and, if needed, use more cores. The example is a report over an access log, written three times.
The project is in examples/apps/perf-lab. The raylang blocks are
copied from it and CI checks that they still are. The figures are from a MacBook Pro M3 Pro, the
median of five runs: on your machine they will differ, the ratios much less.
1. Measure before touching anything
The program generates its own input, always the same, and times only the work being compared.
time.monotonic_millis() is a clock that never goes backwards, the right one for measuring:
// Time only the work being compared, not the generation of the input.
let start = time.monotonic_millis();
let report = if (which == "slow") {
slow.report(lines)
} else if (which == "parallel") {
parallel.report(lines, 4)
} else {
fast.report(lines)
};
let elapsed = time.monotonic_millis() - start;
ray run -- slow 200000
slow: 200000 lines, 6005 paths, 141343 chars, 7274 ms
Seven seconds for 200,000 lines. Before optimizing, you need to know where they go.
2. Where the time goes: ray profile
ray profile runs the program and, when it ends, prints a table per function: self time (without
what it calls), the percentage, the time including its calls, how many times it was called and the
average.
ray profile -- slow 200000
self ms self% incl ms calls avg µs function
5332.172 85.3% 5332.172 1 5332171.67 slow::order
731.924 11.7% 731.924 200000 3.66 position
90.331 1.4% 822.255 1 822255.21 slow::count
77.339 1.2% 77.339 1 77338.88 data::log
17.885 0.3% 17.885 1 17884.71 slow::render
The first version has three steps, each in its own function, and the profile tells them apart:
slow::ordertakes 85%. It is a hand-written selection sort: for each path it walks all the others, 36 million iterations for 6005 paths.positionis 12%. The count keeps the paths in an array and looks each line up by scanning it.slow::renderis 0.3%. Building the text by concatenating in a loop looked like the usual suspect, and it weighs nothing.
That last line is the reason to measure: intuition would have started in the wrong place.
Two tips for reading the profile. Some library functions, such as position, get or sort_by,
show up under their own name; basic operations, such as split, to_string or concatenation, are
counted inside the function that uses them. And a profile only tells functions apart: if all the
work is in one, split it into steps and measure again.
3. Fix the algorithm
This is the count in the first version, with its linear search:
// Requests and total milliseconds per path.
fn count(lines: [string]) -> Tally {
let t = Tally { paths: [], counts: [], totals: [] };
for line in lines {
let parts = line.split(" ");
let path = parts[1];
let ms = parts[3].parse_int().unwrap_or(0);
// Linear search: every line scans the paths seen so far.
match (t.paths.position(path)) {
Option.Some(i) => {
t.counts[i] = t.counts[i] + 1;
t.totals[i] = t.totals[i] + ms;
},
Option.None => {
t.paths.push(path);
t.counts.push(1);
t.totals.push(ms);
},
}
}
t
}
The second version replaces the array with a Map, where finding a path costs the same with ten
or ten thousand of them:
/// Requests and total milliseconds per path.
pub fn tally(lines: [string]) -> Map<string, Stat> {
var stats: Map<string, Stat> = Map.new();
for line in lines {
let parts = line.split(" ");
let path = parts[1];
let ms = parts[3].parse_int().unwrap_or(0);
// Hash lookup: constant time, however many paths there are.
match (stats.get(path)) {
Option.Some(s) => {
s.count = s.count + 1;
s.total = s.total + ms;
},
Option.None => stats.insert(path, Stat { path: path, count: 1, total: ms }),
}
}
stats
}
Structs have reference semantics: s.count = s.count + 1 modifies the value that is in the map,
without inserting it again. The hand-written sort becomes sort_by, and the text is composed by
joining the lines once:
/// The report text: one line per path, most requested first (ties by path).
pub fn render(stats: Map<string, Stat>) -> string {
let sorted = sort_by(stats.values(), fn(a: Stat, b: Stat) -> bool {
a.count > b.count || (a.count == b.count && a.path < b.path)
});
// Collect the lines and join them once.
var out: [string] = [];
for s in sorted {
out.push(s.path + " " + to_string(s.count) + " " + to_string(s.total / s.count));
}
out.join("\n") + "\n"
}
fast: 200000 lines, 6005 paths, 141343 chars, 143 ms
From 7274 ms to 143 ms: 51 times faster, without changing engines. A test checks that both versions produce exactly the same report; without that test, an optimization is a bet.
4. Compile to native
ray run executes the program on the VM. ray build --native translates it to Rust and compiles
it to machine code, with the same output byte for byte.
ray build --native --release -o perf-lab
./perf-lab fast 200000
| 200,000 lines | VM | Native |
|---|---|---|
| First version | 7274 ms | 614 ms |
Version with Map and sort_by |
143 ms | 23 ms |
The native binary is 6 to 12 times faster on this program, and uses half the memory (124 MB against 242 MB with two million lines). But compiling the first version to native leaves it at 614 ms: four times slower than the good version on the VM. The algorithm comes first.
The build options tune what is left:
--releaseturns on every Rust optimization. Here it improves the work with strings and maps by 20%; it takes longer to compile, so it is for the deliverable.--fastswaps checked arithmetic for arithmetic that does not detect overflow. On this program it changes nothing measurable; it helps in numeric loops, at the price of a guarantee.
To compare with other languages, the benchmarks page measures the native binary against Go, Rust and Node on fourteen programs.
5. Use more cores
Every raylang fiber has its own memory: there is no shared data and there are no locks. Launching a fiber per slice of work is direct, with one condition you need to know. The fiber receives a copy of what it captures, and its result is copied back.
/// The same report as `fast.report`, computed by `workers` fibers.
pub fn report(lines: [string], workers: int) -> string {
let chunk = lines.len() / workers + 1;
var tasks: [Task<Map<string, Stat>>] = [];
var from = 0;
while (from < lines.len()) {
let part = lines.slice(from, from + chunk);
tasks.push(spawn(fn() -> Map<string, Stat> { fast.tally(part) }));
from = from + chunk;
}
// Merge: add up each path's counts across the slices.
var merged: Map<string, Stat> = Map.new();
for t in tasks {
for (path, s) in join(t) {
match (merged.get(path)) {
Option.Some(m) => {
m.count = m.count + s.count;
m.total = m.total + s.total;
},
Option.None => merged.insert(path, s),
}
}
}
fast.render(merged)
}
| 2,000,000 lines | VM | Native |
|---|---|---|
| One fiber | 1234 ms | 221 ms |
| Four fibers | 460 ms | 111 ms |
Four fibers give twice the speed on native and 2.7 times on the VM, not four: copying the lines to
each fiber has a cost. With RAYLANG_THREADS=1, which forces a single thread, the parallel version
takes 296 ms: more than the one-fiber version, because it pays for the copies without gaining
cores. Splitting the work pays when each slice costs much more than copying it.
6. Servers
On a web server almost none of the above is needed: every connection already runs in its own fiber and the server uses all the cores. What matters there is something else:
- A connection pool to the database, shared between requests, as in the API chapter. Opening a connection per request is the cost you notice most.
- The native binary. The framework's server, compiled, serves on the order of 188,000 requests per second in the project's load benchmark, with about 21 KB per connection.
app.gzip()for large responses, andstatic_embeddedfor static files.
In short
- Measure with a clock and a fixed input.
- Profile with
ray profileand believe the table, not your intuition. - Fix the algorithm: here, 51 times.
- Compile to native with
--release: another 6 times. - Spread across cores only the work that weighs more than its copy: here, 2 times.
Next step
Shipping: signing, publishing and updating what you have built.
<!-- sync: sha256:e24e1bd798b0 -->