Your degree’s university matters more than your major. A HESA-based analysis of nearly 1.9 million UK graduate records found that the institution attended—referred to as “Provider”—accounts for 14.5% of the variance in starting salaries. This exceeds the 9.4% contribution of a graduate’s age and the 8.1% impact of their academic year. Yet most institutions publish only a single blended “average salary,” obscuring these decisive factors.
The problem deepens as AI search becomes the primary lens for career information. When students ask an AI model why a certain program leads to higher pay, the system relies on structured, attribute-level data to generate a specific answer. A flat average salary figure provides no causal insight, making the institution invisible to AI-driven discovery. How should higher education leaders present grad outcomes to ensure their distinct value propositions are understood by both humans and the AI visibility tools that now shape career decisions?
The 12 variables that actually determine graduate salaries
To understand what truly drives pay, Henshaw et al. analyzed a dataset of 1.87 million records from the Higher Education Statistics Agency (HESA) covering the UK. By applying SHAP (Shapley Additive exPlanations) analysis to a machine learning model, they moved beyond simple averages to identify specific attributes with measurable statistical impact. This explainable AI approach is crucial because a single aggregate number cannot reveal the nuanced factors that differentiate individual career trajectories.
The study highlights that institutional reputation is the single largest determinant of salary. Below, we detail the most significant variables identified in the analysis:
| Rank | Variable | Impact on Salary |
|---|---|---|
| 1 | Provider (Institution) | 14.50% |
| 2 | Age of Qualifier | 9.41% |
| 3 | Academic Year | 8.05% |
| 4 | Post-Graduation Activity | 7.65% |
| 5 | Socio-economic Classification | 6.34% |
| 6 | Domicile | 5.96% |
These specific drivers, such as domicile and socioeconomic classification, are the data points that must be surfaced for AI visibility. Generic aggregates fail to provide the context needed for generative engines to construct accurate, attribute-level answers regarding program value and graduate outcomes.
Why average salary data fails for AI search
A static average is a single number that aggregates thousands of diverse careers into one blunt figure. In contrast, attribute-level outcomes are specific, explainable data points that link individual factors—like major, institution, or location—to concrete earnings results. The difference matters because it determines whether an algorithm can parse the “why” behind a salary, not just the “what.”
Generative AI engines do not read a page to get a vibe. They scan for specific, citable facts to construct answers to complex user queries, such as “what drives a high starting salary in engineering?” When a search engine encounters a flat salary figure, it lacks the granular data needed to attribute cause and effect. Without that causal link, the AI cannot verify which specific program characteristic—such as curriculum depth or industry partnerships—led to higher earnings. The result is generic, low-confidence answers that fail to distinguish one program from another.
Consider the HESA study’s finding that institutional reputation is the single strongest predictor of graduate pay, accounting for 14.5% of the predictive power. If a university publishes only its average salary, an AI engine sees a static value but has no context on why that institution’s graduates earn more. It cannot cite “institutional brand” as a differentiator because the data does not explicitly link that variable to the outcome. Consequently, the AI treats the salary as a coincidence rather than a result of a specific institutional impact. For institutions relying on AI visibility to attract students, this omission effectively erases their competitive advantage from the digital record. The algorithm simply cannot articulate what makes the program valuable if the data doesn’t tell it.
Structuring your program outcomes for AI retrieval
To achieve meaningful AI visibility in education, institutions must move beyond static aggregates toward a structured, attribute-level reporting framework. Mirroring the 21-variable taxonomy from the HESA study provides a solid foundation for this shift, even if your specific institutional variables differ. By adopting this comprehensive framework, you ensure that your grad outcomes are captured with the granularity required for modern search algorithms to interpret and cite.
The technical execution of this strategy is just as critical as the data collection itself. Raw data buried in PDF reports or narrative summaries is inaccessible to generative AI engines. For effective retrieval, metrics must be stored in machine-readable formats such as JSON-LD or clean, relational tables. This structure allows attribute-level metrics to be explicitly linked to specific student outcomes. Without this technical backbone, AI models cannot parse the relationship between a specific variable, such as domicile, and its impact on earnings.
Consider the difference in clarity between a generic statement and a data-rich one. A traditional report might simply state: “Average: $50,000.” A structure optimized for AEO education would instead publish: “Graduates from [Institution] in [Field] with [Credential] have a median of [Amount], with [Factor X] showing a [Y%] impact on pay.” This format provides the specific, citable facts that AI engines look for when answering user queries about career prospects. It transforms a static number into a dynamic, explainable data point that highlights the actual drivers of salary growth.
Metadata plays a pivotal role in ensuring these data points are interpreted correctly by AI models. Terms like “domicile” or “socioeconomic classification” carry specific meanings within the HESA framework that may not be universally understood. Clearly defining these terms in your metadata ensures that AI systems correctly interpret the nuances of the data. If a model misinterprets a term, it may generate inaccurate answers about the value of your program, ultimately undermining your efforts to provide accurate salary data to prospective students and employers.
Frequently asked questions on education data for AI
Does publishing salary data really change AI visibility?
Yes. AI engines prioritize specific, verifiable data points over vague marketing claims when generating answers for students and employers. If you provide granular facts, the model has concrete evidence to cite in its response, making your institution a recognized source of truth in generative search.
What is the most important metric to publish first?
The ‘Provider’ or institutional reputation metric should be your starting point. The HESA study identifies this as the primary driver of salary differentiation, accounting for 14.50% of the predictive power. This variable outweighs even the field of study, making it a critical anchor for explaining why grad outcomes vary significantly across different institutions.
How should I handle sensitive data like socioeconomic background?
Publish aggregated, anonymized trends rather than individual-level data. For example, you can report that graduates from higher socioeconomic classes had average salaries approximately 25.03% higher than those from lower classes. This approach maintains strict privacy standards while still providing the value needed to explain wage gaps in the context of AEO education. By focusing on the impact of these factors on outcomes, you offer insights without compromising student confidentiality.
Is the HESA model applicable to non-UK institutions?
The specific variables are UK-centric, but the methodology is universal. The core approach of identifying ‘attribute-level’ drivers via explainable AI applies to any higher education context. Whether you are in North America or Asia, the principle remains the same: breaking down aggregate salary data into specific, explainable factors allows AI systems to construct accurate, nuanced answers about your program’s value.
Conclusion
The shift from static reporting to dynamic, explainable data publishing is not merely a technical upgrade; it is a strategic necessity for AI visibility. As generative engines become the primary gateway for career information, institutions that publish only aggregate figures risk having their value proposition obscured by the very algorithms they seek to engage. Decisions about higher education are increasingly mediated by AI models that prioritize specific, attribute-level facts over generic marketing claims. If your grad outcomes are not structured to allow these systems to identify true drivers like institutional reputation or socioeconomic factors, the resulting answers will be vague or misleading. The strategic implication is clear: control over how your education data is interpreted in the AI era depends on providing the granular, verifiable context that allows for accurate attribution. The question for decision-makers is no longer whether to publish data, but whether that data will speak clearly in the language of explainable AI.