Pseudorandom Features

To support changes which aren’t fully consistent, tadpole offers a handful of features to add randomness.

The engine uses a deterministic PRNG seeded with a hash of the input word, so that for the same rules and input word the result will be consistent, but changing the input word even slightly will yield different random outcomes. Changing the ruleset may change the number and order of PRNG calls, so if you add or remove randomness elements results may change for words that might not intuitively be affected.

For clarification, whenever this chapter uses the term “random” it is meant to refer to deterministic pseudorandom numbers.

Random Annotation

A rule can be randomly applied to some percentage of words with the random annotation.

rule half_the_time = #{50%} 'a' -> 'b';

This will replace a with b in 50% of words.

Random Replacer

A replacer can be prefixed by a percentage in parentheses to make it apply only sometimes.

rule random_replacer = 'a' -> (50%)'b';

This may seem redundant with the random annotation, but these behave different in compound rules. Random annotations will prevent the rule from generating matches, while random replacers restore the original matched symbols instead of inserting the replacer content. This means that for the purpose of compound rules, a random-annotated rule which doesn’t apply didn’t generate any matches, while a rule using a random replacer did, affecting flow in fallback rules or converging rules.

Since random replacers work by restoring the original symbols, they should in most cases only be applied to the entire replacement and not only to sections of it.

Distributions

Distributions are a special type of binding that can be applied to replacers expecting an index binding. They generate an index randomly based on the given random distribution.

Uniform Distribution
rule uniform_distr = 'a' -> vow@(uniform);

The uniform distribution creates a random index, choosing all possible values with equal likeliness.

Zipf Distribution

Zipf’s law is an observation that, among other things, the second most common word in a large corpus of text is half as common as the most common, the third most common is a third as common as the most common, and so on. This observation applies to a surprising amount of other things as well.

This principle can be generalized into the Zipfian distribution:

𝑃𝑁𝑠(𝑖)=1𝑖𝑠𝐻𝑁𝑠𝐻𝑁𝑠=∑𝑘=1𝑁1𝑘𝑠

Where 𝑁 is the number of values to choose from and 𝑠 is a parameter which can be set by the user.

rule zipf_distr = 'a' -> vow@(zipf 1.0);

This requires the classes or alternatives the distribution to be ordered in the desired frequency order. The parameter 𝑠 can be any non-negative number, with 0 resulting in an uniform distribution and higher values weighting earlier elements more heavily.

For more on the Zipfian distribution, refer to the Wikipedia article on Zipf’s Law.

Explicit Weights

If the above distributions aren’t appropriate, It is also possible to specify weights for each element explicitly. This has the disadvantage of needing to manually write out weights and match counts.

rule weights = 'a' -> ('a' | 'b' | 'c' | 'd')@(1 2 3 4);

The probability for each element in a weighted distribution is

𝑃(𝑖)=𝑤𝑖∑𝑘=1𝑁𝑤𝑘

Where 𝑤𝑖 are the the weights and 𝑁 is the total number of them.