Sequential & Parallel
Flow blocks are the scheduling primitives of Kathryn. You write them as Python
context managers inside a @flow method; every assignment placed inside a block
attaches to that block, and the block decides when — on which clock
cycles — those assignments fire. This page covers the two skeleton blocks that
everything else nests inside: seq and par.
Naming convention
Section titled “Naming convention”Before diving in, learn the prefix convention used across all flow blocks — it tells you how a construct treats the clock:
| Prefix | Meaning | Examples |
|---|---|---|
c- | condition checked combinationally — the check itself costs no extra cycle (the body still takes its own cycles) | cif, cselif, cselse, cwhile, cdowhile |
s- | condition sampled sequentially — one extra clock is spent registering the check | sif, swhile, scwait |
z- | zero-cycle — the whole construct is pure gating logic and consumes no cycles at all | zif, zelif, zelse, zstate, zcase |
p- | pick family — gated multi-way selection | pick, pif, pidef |
Two names sit slightly outside the pattern: cloop is a counter loop, and
sywait is a fixed cycle wait (see Waits).
seq — one operation per cycle
Section titled “seq — one operation per cycle”A seq block runs its contents in order, one step per clock cycle. Each direct
assignment (or nested sub-block) gets its own state in a chain; when a state’s
bit is high, its assignment latches on that clock edge, and the bit hands off to
the next state.
Adapted from tc1_seq_simple:
class tc1_seq_simple(Module): @init def com_declare(self): self.x = reg(8, "x") self.y = reg(8, "y") self.simple_val = val(8, 48, "simple_val")
self.x.mark_output("my_x") self.y.mark_output("my_y")
@flow def my_flow(self): with seq(): self.x |= self.simple_val # cycle 1: x <= 48 self.y |= self.x # cycle 2: y <= x|= is the clocked assignment operator (*= is the combinational one —
see Assignment). Cycle by cycle, after the
master reset mrst is released:
- Edge 1 — the internal
startpulse sets the first sequence state bitseq_state_0. - Edge 2 —
seq_state_0is high, sox <= 48latches; the chain advances (seq_state_1 <= seq_state_0). - Edge 3 —
seq_state_1is high, soy <= xlatches. Becausexalready latched on the previous edge,yreceives 48.
The emitted Verilog makes the state chain explicit. Each assignment is guarded by its own one-bit state register, and each state register is set from the one before it:
always @(posedge WIRE_clk) begin if (SR_ST_seq_state_4_0_ST) begin REG_x[7:0] <= VAL_simple_val[7:0]; // step 1 endend
always @(posedge WIRE_clk) begin if (SR_ST_seq_state_4_1_ST) begin REG_y[7:0] <= REG_x[7:0]; // step 2 endend
always @(posedge WIRE_clk) begin SR_ST_seq_state_4_1_ST <= 1'b0; if (SR_ST_seq_state_4_0_ST) begin SR_ST_seq_state_4_1_ST <= 1'b1; // chain: state 0 -> state 1 endendThe state chain, one step per cycle:
stateDiagram-v2
[*] --> seq_state_0: edge 1 start pulse
seq_state_0 --> seq_state_1: edge 2 latch x = 48
seq_state_1 --> [*]: edge 3 latch y = x
par / par_auto — concurrent operations
Section titled “par / par_auto — concurrent operations”A par_auto block runs its contents at the same time. All direct
assignments inside it are merged under one shared state, so they all latch on
the same clock edge. par is simply an alias for par_auto — both build the
same auto-synchronized parallel block.
Adapted from tc2_par:
@flowdef my_flow(self): with seq(): with par_auto(): self.x |= self.val_5 # both latch self.y |= self.val_10 # on the same edgeCycle by cycle: edge 1 arms the sequence state; edge 2 latches x <= 5 and
y <= 10 together. Contrast with seq, where the same two statements would
take two edges.
Exit synchronization
Section titled “Exit synchronization”When a par block contains nested sub-blocks, its branches may take different
numbers of cycles. The variants differ in how the block decides it is finished
(which matters when the par sits inside a seq that must continue
afterwards):
par_auto— the exit is auto-synchronized. If the branches take the same statically known number of cycles, the block exits with them and no extra hardware is needed; if their lengths differ, the compiler inserts a synchronizer node so the block’s exit fires only after every branch has finished.par_no_sync— no synchronizer is inserted. With branches of differing length the exit is the OR of the branch exits, so the enclosing flow moves on as soon as any branch completes.
How the two variants decide the block’s exit:
flowchart TB
S["par block entry"] --> A["branch A"]
S --> B["branch B (may be slower)"]
A --> J{"exit rule"}
B --> J
J -->|"par_auto: exit after ALL branches finish"| E["exit fires"]
J -->|"par_no_sync: exit after ANY branch (OR)"| E
Writes to the same register in par
Section titled “Writes to the same register in par”Two branches of a par may write the same register in the same cycle. The
result is not a race: Kathryn resolves the writes by priority, and ties are
broken by program order (the later statement wins under non-blocking
semantics). Adapted from tc12/tc13:
with seq(): with par_auto(): with priority(100): self.x |= hi_src # higher priority: emitted last, dominates with priority(50): self.x |= lo_srcSee Write Priority for the full rules.
Nesting
Section titled “Nesting”Skeleton blocks nest freely: a seq step may be a whole par block (which
counts as one step lasting as long as its slowest branch), and a par branch
may be a whole seq chain. The cif/sif conditionals and the loop blocks
(Conditionals, Loops)
automatically open an inner skeleton for you — a seq when nested in a seq,
a parallel skeleton when nested in a par — so their bodies behave like the
enclosing context.