AssessWikiAssessWiki

Policy

How AssessWiki evaluates, rates, and reviews assessment tools.

Last updated: August 1, 2026

In short

AssessWiki evaluates each assessment tool against five criteria: theoretical grounding, construct validity, reliability, accessibility, and transparency. Each criterion is scored on a 0–5 scale. Tools must meet a minimum threshold to be listed. We do not accept payment for reviews, ratings, or placement.

1. How we select assessments for review

We proactively monitor the assessment landscape across four sources:

2. Evaluation criteria

Each assessment is evaluated on five criteria, scored 0–5:

Theoretical grounding (0–5)

Is the assessment based on a recognized psychological framework? We trace each tool to its originating theory (e.g., Big Five, Holland Codes, attachment theory) and verify that the items and scoring align with the published framework.

5 = Directly derived from peer-validated framework with documented alignment. 0 = No identifiable theoretical basis.

Construct validity (0–5)

Does the assessment measure what it claims to measure? We examine factor structure, convergent validity (correlation with established measures), and discriminant validity (non-correlation with unrelated constructs).

5 = Published independent validation studies. 0 = No validity data available.

Reliability (0–5)

Are results consistent? We look for internal consistency (Cronbach's alpha ≥ 0.70), test-retest reliability (r ≥ 0.70 over 2+ weeks), and inter-rater reliability where applicable.

5 = Multiple published reliability studies with strong results. 0 = No reliability data.

Accessibility (0–5)

Is the tool usable by its intended audience? We assess cost (free/paid/freemium), language availability, device compatibility, time to complete, reading level, and accommodations for disability.

5 = Free, multilingual, responsive, WCAG-compliant, <15 min. 0 = Expensive, single-language, desktop-only, >60 min.

Transparency (0–5)

Does the publisher disclose methodology? We check whether scoring algorithms, norming data, limitations, and conflicts of interest are publicly documented.

5 = Full disclosure of methodology, norms, and limitations. 0 = "Black box" scoring with no published methodology.

3. Overall rating

The five criteria are weighted and combined into an overall score (0–25), which maps to a rating tier:

ScoreRatingMeaning
22–25RecommendedStrong evidence across all criteria
18–21SolidGood evidence with minor gaps
13–17LimitedSome evidence, significant caveats
8–12WeakInsufficient evidence, use with caution
0–7Not recommendedLacks theoretical or empirical basis

Tools scoring below 8 are generally not listed in our library. We may list them with a "Not recommended" rating if they have significant user adoption and warrant a public warning.

4. Update and review cycle

5. Conflicts of interest

6. Corrections and feedback

We welcome corrections from assessment authors, researchers, and users. If you believe a review contains an error, please contact us with:

We review all correction requests within 14 days and publish corrections with a dated changelog.