Tuesday, February 28, 2017

An immediate interpretation of probability

Given that many interpretations of probability have been considered and rejected*, it is not clear that a simple interpretation of probability exists. For example, the rhetorical flourishes of the simple interjection “Fat chance!” are unlikely to be captured in a rational* discussion. Furthermore, even if such an interpretation did exist, it is not at all obvious that a paper should be written about it. For example, a study of monkeys, who share 93% of DNA with humans*, shows that they do not use writing in the wild. Monkeys are nonetheless able to learn simple mathematics* and English*. We thus defer discussion of the subject by treating the paper as a preliminary mere-exposure experiment*; it does not need to be approved in advance due to essays falling under normal educational practices*.
Our interpretation may be termed immediacy. Consider the statements “Alex is wearing a red sweater” and “Alex is wearing a blue sweater”, and assume that the former was observed yesterday but the latter is present today. Qualitatively, these perceptions are different. We can interact with the blue sweater, for example by splashing paint on it; we term this “near”. In contrast, the red sweater exists only in our mind; we may have recollected incorrectly*, for example, if Alice was wearing a red sweater but swapped desks with Alex; we term this “far”. The far reality is mentally constructed, with no direct perceptual connection.
Far reality is fragile. Consider a child with no experience of elevators, taking their first ride alone. The door closes on their parents, a number changes, the elevator beeps, and when the door opens again a new space is presented, with complete strangers and no parents in sight. Confusion results, only resolved by introducing a notion of “floor” and its corresponding sensations and notifications of vertical movement.
A reasonable reaction is to ask is how to avoid the possibility of being confused. Unfortunately, we run into the problem of incompleteness*, and in particular the problem of other humans*. Although a problem may not necessarily have a solution*, it may have an approximation; furthermore, we can categorize the techniques used to approximate the solutions, in this paper termed “mental constructions".
What are mental constructions? Again, an example: “The visual system, although highly complex*, in its most basic elements appears to function as a recording and processing device*.” The previous sentence is simply a reflection of the dominance of the computational paradigm*. A better model than simple computation is the act of writing itself, wherein letters are sequentially produced on a page by some complex process with memory*. We of course lose aspects of humanity when we attempt to describe it in writing, but it is futile to expect to solve these problems via more writing*, so we stop there. Thus, for our purposes, mental constructions are recurrent processes that result in the production of writing. Finally, then, we arrive at our interpretation of probability; given several possible processes with which to produce writing, select one (for this paper, a randomized variant of the autofocus system*).
One particular strength of our interpretation is its similarity to the etymology of probability*. The word probability originates from the Proto-Indo-European root *pro-bʰwo-, meaning “‎to in front”. This came to be associated with the Latin probus, meaning good, noble, and virtuous. After morality became a common good*, there were then techniques of probo to test and certify what was good. Hypothetically this then became the Latin probābilis, meaning provable, credible, and finally probability as a measure of provability. In our case, provability is determined by a measure of verbosity; in particular, we measure proof as an estimate of the amount of processing done, via an analysis of the amount of references and obscure concepts used.
Another strength is its self-referential nature, more precisely homiconicity*; by writing a paper interpreting probability using a probabilistic paper-writing process, we achieve a similar representation of both paper and probability. Our interpretation thus covers a vast conceptual space in a small amount of time by utilizing foundational tools repeatedly, particularly the Axiom of Choice*. Although this might appear to be a weakness, in that the two are mutually recursive and thus could lead to an infinite regress, there are well-established base cases of an empty paper and not writing a paper. Furthermore, the homoiconic property ensures that the presented concepts can be reproduced by reproducing the paper.
A final strength of our interpretation is its synthesis of concepts present in frequentist and Bayesian probability. By its emphasis on writing, we mimic the emphasis on enumeration of possibilities found in frequentism. However, whereas a frequentist approach might focus on the space of possibilities, e.g. H and T in a coin flip, we fix the choice of possibility space (a sequence of English characters) and instead focus on the choice of process. Furthermore, we leave the general method of choice purposely ambiguous, thus implicitly introducing subjectivist notions of probability. But unlike Bayesian probability we successfully avoid the need to enumerate all possible choices of process, by introducing a stopping procedure (detecting the presence of self-reference). Our interpretation of probability is thus strictly constructionist* and does not introduce problems of computability or practicality.
Our interpretation does have a flaw, though, in that it does not guarantee optimality in any sense. For example, it does not preclude the possibility of accepting a Dutch book. Although it heavily emphasizes writing, references, and consideration of the means of production*, it addresses the possibility of an unconsidered approach only reactively. Furthermore, it does not necessarily produce agreement among different individuals*; they may make different choices in process. We submit three responses: first, that optimality is a theoretical concern* of no particular relevance to applied mathematics*; second, that our approach is optimal in terms of a suitably defined global measure of cognitive costs*; and third, that one can always postulate an independent third party that would use the process of considering and synthesizing all other processes*, but such a third party does not necessarily exist in practice.
Although any experiment is subject to uncertainty, we have attempted to minimize it through the use of a controlled environment and hypothetical situations. The most pressing concern is that the experiment could fail to produce any measurable effect. Particularly, its focus on references and concepts makes it conceptually difficult to understand, and thus it may be rejected altogether. Furthermore, the use of randomized concepts may make transitions in the paper appear stilted or disjointed. Time constraints forced the summarization, omission, and consequent misunderstanding of many concepts that require an interpreting paper in their own right. Nonetheless we look forward to seeing the results.

Wednesday, May 11, 2016

Fedora workflow diagram

workflow buildbot copr/koji builder cached_builds build cache (RPM) buildbot->cached_builds autosigner bodhi bodhi buildbot->bodhi pkg list stable stable mirrors mirrors stable->mirrors admins sync bug_fixed bug fixed - QA examines stable->bug_fixed release release stable->release users users users->bodhi feedback bug_untriaged bug_untriaged users->bug_untriaged new issue or automated crash report mq local repo patch patch/tests mq->patch MAR rpm-ostree cached_builds->MAR repo repo cached_builds->repo taskotron taskotron cached_builds->taskotron tests try run testsuite patch->try incoming git repos / pkgdb patch->incoming hotfix MAR->mirrors tbpl build/test results by changeset (resultdb) tbpl->incoming backout regressions bug_new bug_new tbpl->bug_new results tbpl->bodhi auto-updates try->buildbot scratch reviewer review try->reviewer r? incoming->buildbot CI incoming->mq rebase bug_assigned bug_assigned incoming->bug_assigned backout testing testing testing->stable promoted testing->users power frozen frozen/beta testing->frozen release driver frozen->stable mirrors->users common usage mirrors->mq packager specfile specfile mirrors->specfile anitya & friends specfile->buildbot scratch build specfile->bug_new submit package bug_new->bug_assigned development bug_useless unusable bug bug_new->bug_useless triage repo->mirrors modules ISO ISO repo->ISO pungi ISO->mirrors bodhi->buildbot tags bodhi->stable approval bodhi->testing admins sign bodhi->taskotron test taskotron->tbpl bug_fixed->bug_new reopened bug_verified bug_verified bug_fixed->bug_verified wiki relnotes bug_fixed->wiki release->mirrors bug_untriaged->bug_new module owner / popular vote bug_untriaged->bug_untriaged resolved duplicate bug_assigned->mq new branch bug_assigned->bug_new nevermind reviewer->incoming push approved reviewer->bug_assigned r- bug_useless->bug_verified officially unsupported bug_closed bug_closed bug_verified->bug_closed Bug is not definitely needed anymore

Wednesday, May 20, 2015

Mnemosyne data analysis

I did this a while ago.

Steps:

  1. Download 2014-01-27-mnemosynelogs-all.db.xz
  2. Extract: tar -xvf 2014-01-27-mnemosynelogs-all.db.xz
  3. Count the records: sqlite3 -batch "select event, count(*) From log Group by event;"

    Total121 188 408Event type
    1 813 548start
    2 786 586stop
    3 172 462scheduler
    4 3 123 878load db
    5 1 109 352save db
    658 532 022add card
    8 7 684 724delete card
    948 965 836repetition
  4. Dump repetitions to CSV: sqlite3 -csv -header -batch 2014-01-27-mnemosynelogs-all.db "select object_id,grade,acq_reps,ret_reps,actual_interval From log where event=9 limit 7;" (this only does the first few lines of course)

    c136315a,9779a1ad,1,0,0,5
    c136315a,a2a80b21,1,0,0,4
    c136315a,a35a6c4a,1,0,0,5
    c136315a,85f8ec88,1,0,0,5
    c136315a,10ae2adc,1,0,0,5
    c136315a,a9c66681,1,0,0,4
    c136315a,ba841422,1,0,0,5
    c136315a,4e108d46,1,0,0,5
    
  5. Sort on time to completion and grade:
    db9864c5,37132b6b,6,10,-188957807,3
    db9864c5,680c6178,3,7,-188879971,3
    db9864c5,40712c82,7,6,-188879963,3
    db9864c5,a5679e62,17,14,-188879941,3
    db9864c5,21a9326f,14,11,-188652473,3
    dQjTu8hqz3d04ReDrWTxdZ,7humHviBhzGDsbmfZzkFNh,2,0,-157680000,3
    dQjTu8hqz3d04ReDrWTxdZ,LGuN7QIDz7jsjuwKeSazrB,2,0,-157680000,3
    dQjTu8hqz3d04ReDrWTxdZ,a2cgcETy4U1mpR8gId581B,2,0,-157680000,2
    dQjTu8hqz3d04ReDrWTxdZ,2dHnVSa6v7HJXd1Bk6oOI7,2,0,-157680000,0
    dQjTu8hqz3d04ReDrWTxdZ,evdCglXJdRSwNAJwOaxDJo,2,0,-157680000,0
    dQjTu8hqz3d04ReDrWTxdZ,iWboGsLywTzw2agIQPUsBw,3,1,-157680000,0
    dQjTu8hqz3d04ReDrWTxdZ,y4hyxaB2HdW6UVQn15vdoh,23,2,-157680000,0
    qFHzllJkeSgoONQZJxkJ3c,LWYCS7wgH2T3aHPRUVUnH4,3,0,-127916114,2
    qFHzllJkeSgoONQZJxkJ3c,CIbwahRHaTaTCujq1X10QB,3,0,-127916114,0
    qFHzllJkeSgoONQZJxkJ3c,QPuSfCkJzbzk7uw2AJEA73,5,0,-127916103,1
    qFHzllJkeSgoONQZJxkJ3c,zdpemIGjakRrDMmK7Lvc6d,2,0,-127916095,1
    qFHzllJkeSgoONQZJxkJ3c,ILfYUuLjwER4VnpIYtUpWj,6,0,-127916088,0
    qFHzllJkeSgoONQZJxkJ3c,GdkgWUvUZMkhbp4QZim6fA,13,0,-127916071,1
    
  6. Run this Haskell:
    {-# LANGUAGE ScopedTypeVariables #-}
    
    module GradeMunge where
    
    import qualified Data.ByteString.Lazy as BL
    import Data.Csv.Streaming
    import Data.Csv(encode)
    import System.IO
    import qualified Data.Map.Strict as M
    
    type Data = (Integer,M.Map Integer Integer,(Integer,Integer))
    
    def :: Integer -> Data
    def r = (0,M.fromList [(0,0),(1,0),(2,0),(3,0),(4,0),(5,0)],(r,r))
    
    main :: IO ()
    main = do
        csvData <- BL.readFile "grades.csv"
        let out = munge (decode NoHeader csvData) (def (-188957808))
        BL.writeFile "gradehisto.csv" (encode out)
    
    munge :: Records (BL.ByteString, BL.ByteString, Integer, Integer, Integer, Integer) -> Data -> [[Integer]]
    munge (Cons (Right (user,obj,acq_reps,ret_reps,interval,grade)) k) d@(total,_,(low,high)) =
        if (total < 10000 || (low <= interval && interval <= high && total < 100000)) then
            -- add the record
            munge k (insertRecord interval grade d)
        else
            -- return the record and start a new one
            prepOut d : munge k (insertRecord interval grade (def interval))
    munge (Cons (Left err) k) a = error ("blah: " ++ err) -- munge k a
    munge (Nil (Just s) k') a = error ("nil: " ++ s) -- [a]
    munge (Nil Nothing k') a = [prepOut a]
    
    prepOut :: Data -> [Integer]
    prepOut d@(total,xs,(low,high)) | length (M.keys xs) == 6 = low : high : total : M.elems xs
                                  | otherwise = error (show d ++ " keys: " ++ show (M.keys xs))
    
    insertRecord :: Integer -> Integer -> Data -> Data
    insertRecord interval grade (total,xs,(low,high)) = (total+1,M.insertWith (+) grade 1 xs,(low `min` interval,high `max` interval))
    
  7. Plot in Mathematica:
And my notes on the database format:
filename: (user_id)_[(machine_id)_](log_number).txt
CREATE TABLE parsed_logs(log_name text primary key); # used for incremental processing
CREATE TABLE _cards(id text primary key,last_rep int,offset int); # used for version < 2 munging; 'last_rep' is the timestamp of the card's last grading time; offset is 0, 1 (grade>=2 phase 1), or -1 (grade <= 2 phase 2)
insert or replace into _cards(id=card_id + user_id, offset, last_rep=)

CREATE TABLE log(
        user_id text,
        event integer,
        timestamp integer,
        object_id text,
        grade integer,
        easiness real,
        acq_reps integer,
        ret_reps integer,
        lapses integer,
        acq_reps_since_lapse integer,
        ret_reps_since_lapse integer,
        scheduled_interval integer,
        actual_interval integer,
        thinking_time integer,
        next_rep integer
    );
# Program Started (program_name_version = Mnemosyne 1.0-RC nt win32)
insert into log(user_id, event=STARTED_PROGRAM=1, timestamp, object_id=program_name_version)
# Program stopped
insert into log(user_id, event=STOPPED_PROGRAM=2, timestamp)
# Scheduler SM2 Mnemosyne
insert into log(user_id, event=STARTED_SCHEDULER=3, timestamp, object_id=scheduler_name)
# Loaded database N N N
insert into log(user_id, event=LOADED_DATABASE=4, timestamp, object_id=machine_id, acq_reps=scheduled_count, ret_reps=non_memorised_count, lapses=active_count)
# Saved database N N N
insert into log(user_id, event=SAVED_DATABASE=5, timestamp, object_id=machine_id, acq_reps=scheduled_count, ret_reps=non_memorised_count, lapses=active_count)
# New item id grade new_interval (munged, possibly add repetition too)
# Imported item id grade ret_reps last_rep next_rep interval (not munged)
insert into log(user_id, event=ADDED_CARD=6, timestamp, object_id=card_id)
# Deleted item id
insert into log(user_id, event=DELETED_CARD=8, timestamp, object_id=card_id)
# R id grade easiness | acq_reps ret_reps lapses acq_reps_since_lapse ret_reps_since_lapse | scheduled_interval actual_interval | new_interval noise | thinking_time
# R id grade 2.5 | 1 0 0 1 0 | 0 0 | new_interval 0 | 0 when adding new card and grade >= 2
insert into log(user_id, event=REPETITION=9, timestamp, object_id=card_id, grade, easiness, acq_reps, ret_reps, lapses, acq_reps_since_lapse, ret_reps_since_lapse, scheduled_interval, actual_interval=timestamp-previous_rep_timestamp, thinking_time, next_rep=timestamp + new_interval)

initial grading in 'Add Card' counted as an acquisition repetition (and explicitly logged as an 'R' event) when grade 2+, otherwise grade= -1 like imported cards

card_id is created as a hash of the card data
gr= the grade, 0-5, default is 0, -1 means "unseen"
easyiness is the easiness parameter from the SM2 algorithm
acquisition reps = # w/ gr<2, including card
retention reps = # w/ gr=2+, including card
lapses = the number of times you forget this card (new grade < 2, old grade >=2)
ac_rp_l, rt_rp_l = ... since lapse
timestamp = last_rep = last actual repetition
next_rep = next scheduled repetition, timestamp
sch_i: scheduled previous interval in seconds
act_i: actual previous interval in seconds
th_t: thinking time in seconds
initial grade 0,1 (failing)   initial grade 2,3,4,5 (passing)
Version < 0.9.8 (phase 1) acq_reps=0    acq_reps=0, R added, acq++
0.9.8 <= version < 2.0 (phase 2) acq_reps=1, but all acq–    acq_reps=1->0, R added
2.0 <= version acq_reps=0    acq_reps=0, R

Monday, July 28, 2014

Workflow diagram

So apparently the web has moved on to HTML5; let's get some diagrams going.
workflow developer developer bug_assigned bug_assigned developer->bug_assigned takes possession mq local repo patch patch/tests mq->patch sheriff sheriff incoming incoming sheriff->incoming tree closure/backout bug_new bug_new / bug_reopened bug_new->developer looked at bug_useless unusable bug bug_new->bug_useless triage module_owner module owner module_owner->bug_new confirmed module_owner->bug_useless buildbot builder cached_builds build cache buildbot->cached_builds testslaves tests tbpl build/test results by changeset testslaves->tbpl test results users users bug_untriaged bug_untriaged users->bug_untriaged new issue or automated crash report cached_builds->testslaves try run testsuite patch->try incoming->buildbot experimental experimental branch incoming->experimental unstable unstable / master incoming->unstable testing testing incoming->testing approval stable stable incoming->stable approval incoming->bug_assigned backout experimental->developer unstable->developer unstable->mq rebase unstable->testing bug_fixed bug fixed - QA examines unstable->bug_fixed frozen frozen/beta testing->frozen release driver frozen->stable stable->users common usage tbpl->sheriff find regressions tbpl->module_owner try->buildbot reviewer review try->reviewer r? bug_untriaged->bug_new popular vote bug_untriaged->module_owner bug_untriaged->bug_untriaged resolved duplicate bug_assigned->mq new branch bug_assigned->bug_new nevermind reviewer->incoming push reviewer->bug_assigned r- bug_verified bug_verified bug_useless->bug_verified officially unsupported bug_closed bug_closed bug_verified->bug_closed Bug is not definitely needed anymore bug_fixed->bug_new bug_fixed->bug_verified wiki wiki documentation bug_fixed->wiki

Thursday, July 24, 2014

Decisions

On unfamiliar ground, such as a changing industry or a new product or business launch, models can be dangerous. In the messy world we live in, they often make invalid assumptions or identify non-existent patterns. Few research efforts have focused on analyzing historical data or identifying similarities between phenomena (which can be considered replications of statistical tools). Thus, most statistical methods are untested and, at best, partial solutions, as the relevant variables and their relationships with outcomes often have no consensus. Although complex regression models have been developed, and shown to be somewhat accurate in their domain, they generally only consider a small subset of the factors that predict success. Even once a statistical method is proven and applicable, many challenges remain unsolved in visualization and user interaction. Interactively and quickly exploring algorithms, parameters, metadata, and data sets requires a large amount of functionality, including zooming, highlighting, filtering, clipping, sorting, smoothing, searching, plotting, focusing, lensing, and side-by-side visualizations. Few graphical toolkit libraries implement all of these or in a fashion that scales to large data sets.

Given this, many decisions still require humans for their consideration and conclusions. Unfortunately, humans are subject to many biases. When presented with a data sample, people naturally begin performing a comparison of the problem with the sample. Due to anchoring, even if the sample is irrelevant, it will still influence their decision. Due to effects such as backfires, bandwagons, decoys, and focusing, it can even influence their answer in the wrong direction. Supposing the data does have a deep correspondence with the problem at hand and creates the correct conclusion, there is still hindsight bias in computing the uncertainty of results and hyperbolic discounting of the payoffs of results.

Using multiple data samples encourages a statistical and historical view of the problem and thus produces more accurate forecasts. It compresses the natural try-fail cycle people use to formulate solutions, allowing people to learn from the past and determine relationships between determinants and outcomes. On average, using more data samples generates more strategies and ensures that the chosen strategy is more successful.

The similarity between samples and the problem has an interesting effect. When generating strategies, it is most useful to have a set of distantly-related or even unrelated samples to consider, as these generate better ideas. However, closely-related examples are a better basis for predicting performance or evaluating strategies, and can greatly reduce over-optimistic evaluations.

It is thus necessary to have both distantly- and closely-related samples for a creative and correct decision. Unfortunately, even when encouraged to do so, people are not naturally inclined to form a broad set of samples; they suffer recall or availability bias. Using a sample set selected randomly from a representative reference class by a unbiased method such as partitioning into intervals or deciles generally reduces this bias. Focusing on the relative similarity of samples and the aspects of them that are generalizable to multiple cases shifts the discussion towards the empirical facts and deep structures of the problem at hand, avoiding superficial details and their unwanted effects. Robust analogizing with differential attention to the most similar and least similar cases increases the precision of outcome analysis and greatly facilitates learning from other people's experiences.

The following procedure incorporates these ideas:
  1. Define the problem and your purpose
    1. predictors, relevant categories or potentialities
    2. the plan/timeline/budget/milestone variables desired
    3. What kind of comparison is being made
  2. Generate a reference class of similar/analogous problems/cases, using as many analogies and references as possible
    1. The class should be universal ("all possible states/choices/outcomes/things of the world")
    2. Common analogies include position, diversity, resources, cycles, economics, and ecology
    3. Make design choices about the study - data sources etc.
  3. Select a set of samples from the reference class. The set should be unbiased and distribution-based, e.g. a "random subset".
    1. predetermine the method of selection, using a rigorously structured approach
    2. avoid memory sampling
    3. avoid using a large set of samples - subsampling allows you to either measure the fragility of a conclusion (if the procedure is repeated) or reduce labor (if not)
  4. Assess the source cases - research
    1. strategies that were pursued, relevant qualities/quantities, and how they turned out ("results")
    2. interaction effects between states, choices, and outcomes - historical analogies
  5. Assess the similarity of the source cases to the target
    1. subjective weighting
      1. rank & rate for similarity - crowdsourcing survey, multiple observers, estimate observer reliability
      2. aggregate with robust mean function - give small weight to outliers
      3. can use non-experts or experts depending on domain / availability of info
    2. rule-based similarity - if relevant features/importance are known
  6. Construct an estimate of results, using the similarity weightings
    1. create measure for comparison of outcome, e.g. money
    2. obtain priors over the probability of the outcome occurring
    3. hierarchical cluster analysis - avg. of ~6 - identify "top" cluster containing samples most similar to problem
    4. estimate average outcomes of samples - use to predict problem
  7. Assess the predictability of outcomes - regression on samples using estimate from (6) and other known variables
    1. Make statistical corrections of the estimate
    2. Adapt or translate results