rulest - GPU Rules Extractor
-
The built-in seeds in
rulestare now five categories of numeric rule chains (prepend/append, mixed, transform+digit, date patterns) automatically generated and tested against the bloom filter in Phase S. These seeds help extract common numeric transformations (e.g., adding years, digits) without requiring manual input.The
--no-builtin-seedsparametr disables Phase S entirely. Use it when:- Your target wordlist contains few or no numeric patterns.
- You want to reduce GPU runtime by skipping thousands of numeric seeds.
- You rely solely on atomic rules (Phase 1) and random chains (Phase 2) or your own
--seed-rulesfile.
Without this flag, Phase S always runs, testing chains up to depth 4 (e.g., ^1 ^2, $1 $9 $9 $0, u $1 and date patterns like $0 $1 $0 $1 $2 $0 $2 $4).
-
New category rules added in the build-in generator
Phase S - Built-in Seed Families (A–M) Numeric families A Pure Prepend digits (depths 1–4) B Pure Append digits (depths 1–4) C Mixed Prepend/Append digits (depths 1–4) D Transform + digit/bracket (depths 2–4) E Date patterns DDMM/YYYY/… (depths 4–9) Special-character families F Pure Append special chars (depths 1–3, top-15 chars) G Pure Prepend special chars (depths 1–3, top-15 chars) H Transform + special char (depths 2–3, top-15 chars) I Digit(s) + special char (depths 2–4, core-7 chars) — covers the ubiquitous "word123!" / "!word123" patterns New families J Leet substitutions (depths 1–2, 10 core pairs) — sa@ se3 so0 si1 sl1 ss5 ss$ st7 sa4 si! — depth 2: leet + digit/special suffix/prefix — depth 2: double-leet chains (e.g. "sa@ so0" → "p@ssw0rd") K Double-transform chains (depth 2, all 15×15 pairs) — covers "c r", "u d", "t f", "E l", "c {", "l ]", etc. L Special-before-digit patterns (depths 2–3, core-7 chars) — reverse orientation of Family I: "!1word" / "word!12" — append: $sp $d… prepend: ^d… ^sp M Leet + transform chains (depth 2) — leet substitution followed by a transform op (all 15) — and transform op followed by leet substitution — covers "P@ssword", "@DMIN", "p@SSW0RD" patterns Special chars - top-15 (F/G/H): ! @ # $ % ^ & * ? . -_ + ( ) Special chars - core-7 (I/L): ! @ # $ % * ? -
Added Phase 3 GA (genetic algotithm) that tries to find good rule chains by mimicking evolution. It starts with a mix of promising chains (from Phase 1 hits) and random ones. Each chain gets a fitness score = how many base words it successfully transforms into target-like words (checked via GPU). Top chains survive as "elite". Pairs of chains swap parts (crossover) and get small random changes (mutations) to create new chains. Weak chains die out. Over many generations, the population shifts toward high-scoring chains. This finds useful rules that random sampling would likely miss. Finally, all discovered chains are merged and deduplicated before output.
An optional evolutionary search that runs after Phase 2 and complements random chain sampling with guided, coverage-driven optimisation. Why it fits this project ──────────────────────── • The fitness function (bloom-filter hits) is already computed by the existing GPU chain kernel — no new GPU code is required. • Phase 2 samples chains *uniformly at random* from the atomic-rule pool. For depth ≥ 3 the search space is |pool|^depth (millions of candidates); the GA focuses probability mass on high-hit-rate regions of that space. • Hot atomic rules from Phase 1 seed the initial population, giving the GA a strong head start rather than searching from scratch. • All Phase-3 discoveries are merged into the global hit counter before signature-based minimisation, so they benefit from the same deduplication and sorting as Phase 1 and Phase 2 results. Algorithm summary ───────────────── 1. Initial population — 30 % depth-2 hot-rule combos, 30 % seeded deeper chains, 40 % random — ensures both exploitation and exploration. 2. GPU-batch fitness evaluation — reuses _run_chain_kernel unchanged. 3. Tournament selection (k = 4) — low-pressure, maintains diversity. 4. One-point crossover (p = 0.80) — exchanges rule-token sub-sequences. 5. Mutation — replace / insert / delete one rule token (weights 60/20/20). 6. Elitism — top <elite_frac> individuals survive unchanged each generation. 7. Diversity guard — duplicate individuals are replaced by random chains. 8. Terminates when <--genetic-generations> is reached or the wall-clock time budget (--target-hours remainder) is exhausted. CLI flags ───────── --genetic Enable Phase 3 (default: disabled) --genetic-generations N Max generations (default: 50) --genetic-pop N Population size (default: 200) --genetic-elite F Elite fraction, e.g. 0.15 (default: 0.15)``` -
Phase 3 GA improvements
-
Novelty‑weighted fitness – Chains not found by Phase 1/S/2 get a 2× bonus during selection → drives GA toward new rules, not rediscovered ones.
-
Unexplored‑seed initial population – The 40 % Phase‑S fill slot prefers chains absent from known_rules; known seeds only as fallback. Depth‑3+ chains biased (70 %) when --max-depth ≥ 3.
-
Dedicated time reservation – 20 % of --target-hours (min 120 s) reserved for GA before Phase 2 starts → GA always gets meaningful runtime.
-
Stagnation guard – No fitness improvement for 5 generations → bottom 30 % of population replaced with fresh random chains (depth‑3+ biased).
-
Depth‑2 warning – Warns that Phase 2 already covers depth‑2 exhaustively → GA adds nothing at depth 2; recommends --max-depth ≥ 3.
-
Functional‑signature registry – Tracks equivalence classes via built‑in probe set. Novelty bonus is functionally aware – equivalent variants get no bonus.
-
Adaptive mutation – If mutated chain falls into a known signature class, up to 2 extra escape mutations applied to break out.
-
Signature‑based offspring filter – Offspring still covered after adaptive mutation → replaced by a fresh random chain (depth‑3+ biased).
-
Honest raw‑hit merging – Stores raw (un‑bonused) hit counts for output; novelty bonus only affects selection.
-
Genuinely novel reporting – Log shows both total GA hits and how many were truly new (absent from known_rules at GA start).
-
-
The Bloom filter — previously one of the slowest parts of tool startup — is now built on the GPU in a seconds, regardless of how large the target wordlist is. The CPU used to grind through it word by word; the GPU does it all at once. If your hardware doesn't support it, it silently falls back to the old way.
-
Benchmark comparison (same corpus)
Single-run results on base = hashmob.mini, target = hashmob.medium
(time-to-result comparison of obtaining similar rule counts: hcrt.pages.dev/debug_rules_efficiency
Full benchmark table: Google Sheet)Ruleset Cover Size Eff. rulest.greedy.150000.rule 36.12% 150000 0.2408 rulest.freq.150000.rule 28.02% 150000 0.1868 rulest.greedy.50000.rule 29.78% 50000 0.5956 rulest.freq.50000.rule 20.25% 50000 0.4050 rulest.greedy.25000.rule 24.91% 25000 0.9964 rulest.freq.25000.rule 15.99% 25000 0.6396 rulest.greedy.10000.rule 19.71% 10000 1.971 rulest.freq.10000.rule 11.44% 10000 1.144 rulest.greedy.1500.rule 10.45% 1500 6.9667 rulest.freq.1500.rule 4.61% 1500 3.0733 rulest.greedy.250.rule 5.06% 250 20.24 rulest.freq.250.rule 1.73% 250 6.92 rulest.greedy.64.rule 2.10% 64 32.8125 rulest.freq.64.rule 0.88% 64 13.75 -
Early benchmark of
rulest_heap*rules after tool update.Ruleset Cover Count Eff. rulest_heap.150000.rule 37.96% 150,001 0.2531 rulest.greedy.150000.rule 36.12% 150,000 0.2408 rulest.freq.150000.rule 28.02% 150,000 0.1868 rulest_heap.50000.rule 33.02% 50,001 0.6604 rulest.greedy.50000.rule 29.78% 50,000 0.5956 rulest.freq.50000.rule 20.25% 50,000 0.4050 rulest_heap.25000.rule 29.59% 25,001 1.1836 rulest.greedy.25000.rule 24.91% 25,000 0.9964 rulest.freq.25000.rule 15.99% 25,000 0.6396 rulest_heap.10000.rule 24.72% 10,001 2.4718 rulest.greedy.10000.rule 19.71% 10,000 1.9710 rulest.freq.10000.rule 11.44% 10,000 1.1440 rulest_heap.5000.rule 20.98% 5,001 4.1952 rulest_heap.1500.rule 15.46% 1,501 10.2998 rulest.greedy.1500.rule 10.45% 1,500 6.9667 rulest.freq.1500.rule 4.61% 1,500 3.0733 rulest_heap.250.rule 9.00% 251 35.8566 rulest.greedy.250.rule 5.06% 250 20.2400 rulest.freq.250.rule 1.73% 250 6.9200 rulest_heap.64.rule 4.77% 65 73.3846 rulest.greedy.64.rule 2.10% 64 32.8125 They fall a bit short of the better ones, but not bad considering the rules were obtained in just one hour and semi-generated. The dictionaries I used were
hashmob.mini(base) andhashmob.medium(target), with all settings enabled:Ruleset Cover Count Eff. concentrator_MT_50000.rule 36.71% 50,000 0.7342 ORTRTS.rule 35.80% 48,439 0.7391 rulest_heap.50000.rule 33.02% 50,001 0.6604 concentrator_MT_25000.rule 31.82% 25,000 1.2728 rulest_heap.25000.rule 29.59% 25,001 1.1836 concentrator_MT_5000.rule 23.38% 5,000 4.6760 rulest_heap.5000.rule 20.98% 5,001 4.1952 concentrator_MT_64.rule 4.92% 64 76.8750 rulest_heap.64.rule 4.77% 65 73.3846 concentrator_MT_64.skull.optimized.rule 4.28% 64 66.8750 - --token-strip
- --token-strip-max-prefix=8
- --token-strip-max-suffix=8
- --genetic
- --genetic-generations=400
- --genetic-pop=800
- --target-hours=0.5 (default)
- --select-mode=greedy
The remaining stages were enabled by default.
-
rulest v3.2– GPU-Resident CELFMajor update focused on performance and memory:
- Full GPU-resident CELF (no more coverage matrix)
- Memory for 14.37M targets: ~6.85 GiB -> ~114 MiB (~61x less)
- Adaptive batch size based on VRAM
- Much faster: (RTX 3060 Ti 8 GB, ~132k base, ~14.37M targets), CELF setup (heap build) - ~40 min -> ~1 min, full selection (150k rules, ~700k candidates): ~15 min instead of ~40 min
Local-search removed. Output format unchanged.
Hello! It looks like you're interested in this conversation, but you don't have an account yet.
Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.
With your input, this post could be even better 💗
Register Login