The raw material of AI: why information architecture decides your university's visibility
Ontology, taxonomies and OKF: why information architecture —not marketing— decides whether AI cites your university or leaves it out.

Key Takeaways
- Content is the raw material: AI doesn't generate knowledge from nothing, it reformulates what already exists. Poorly structured content is the ceiling on everything AI can do with it.
- From ontology to Machine Experience: four linked layers —ontology, taxonomies, content architecture and machine experience— decide whether an agent represents your institution accurately or leaves it out of the conversation.
- A strategic decision, not a technical one: defining what a program is and how it relates to everything else isn't an IT weekend job. It's a structural advantage that's hard to copy.
When a university buys access to an artificial intelligence model, what it really buys is processing capacity. What that model can do with that capacity depends entirely on what you give it to process. AI doesn’t generate knowledge from nothing; it extracts, reorganizes and reformulates it from the knowledge that already exists and is available to it. In that sense, artificial intelligence is an extraordinarily powerful machine that, without the right raw material, produces mediocre or outright wrong results.
The raw material is content. And the way that content is structured —how it’s organized, how it’s classified, how its parts relate to one another— determines the ceiling on everything AI can do inside an institution.
This is why information architecture, a discipline treated for decades as a minor web usability problem, has become a first-order strategic decision. And it’s also why universities that haven’t yet built that infrastructure will find that their AI investments deliver far less than they expected.
The origin of everything: the ontology
Before an architecture exists, an ontology exists. The term sounds more philosophical than it is in practice: an ontology, in the context of knowledge management, is simply the definition of which entities exist in a domain and what kinds of relationships are possible between them.
For a university, the ontology answers questions like: what is an academic program? What distinguishes it from a course, a module, a specialization? What relationship exists between a program and the faculty that teaches it? And between that faculty and the campus where it operates? Can a program belong to more than one field of knowledge? What relationship does a professor have with a program, and is that relationship different if they are the program director or a guest lecturer?
These questions aren’t philosophical. They’re operational. An artificial intelligence system that isn’t clear on what an academic program is —how it differs from other concepts, what attributes define it, what relationships it can have— will make mistakes when talking about it, recommending it or answering questions about it. The quality of the ontology is the lower bound on the quality of the agent.
From the ontology descend the taxonomies: the classification systems that order entities into coherent categories. If the ontology defines what a graduate program is, the taxonomy defines how it’s classified within the institution’s entire academic offering: by subject area, by format, by level, by duration, by target audience. Taxonomies are the controlled vocabulary with which the institution’s various systems —and AI agents— can converse without misunderstandings.
And on top of that ontology-taxonomy pair, content architecture is built: the concrete structure of pages, documents, program records and the relationships between them that shapes the university’s digital presence.
How a graduate program is bought in 2026
To make this concrete, it helps to follow the real process by which a candidate evaluates and decides to enroll in an online university program today.
The search no longer starts on Google and ends on the institution’s website. It starts with a conversational agent —whether an AI-integrated search engine, a chatbot specialized in educational guidance, or a direct query to a language model— and in that first interaction a shortlist of options is already being built. According to data from the consultancy EAB with a sample of more than 5,000 students, 46% already use AI in their process of searching for university programs, and 18% have ruled out an institution directly based on the information an AI system returned to them.
The candidate asks: “Which online digital-marketing master’s of less than a year has the best career outcomes for someone with five years of experience in communications?” The agent doesn’t return a list of links. It returns a synthesis, with names, with compared features, with references to employment rates. Some institutions appear in that synthesis. Others don’t. The difference is not, in most cases, the real quality of the program. It’s the quality of the structure with which the institution has organized and published the knowledge about that program.
Can an AI agent accurately identify the program’s duration? Its format? And distinguish it from the rest of the same institution’s programs? Can it relate the program to its graduates’ employment data? Can it cross-reference that information with the recommended entry profile and the stated career paths?
If the answer is no —because that information exists as free text, scattered across different pages, with no metadata to structure it nor declared relationships between the concepts— the agent will produce an incomplete, imprecise or outright incorrect answer. And the institution will be left out of the conversation at the very moment the candidate needs it most.
The hyperlink and its limit
For thirty years, the hyperlink was the fundamental connection mechanism on the web. And it was a revolution: it turned the isolated document into a node within a navigable knowledge network. The ability to link a program to the academic director’s profile, to the curriculum, to alumni testimonials, to the admissions process, built a user experience that was previously impossible.
But the hyperlink has a structural limit: it doesn’t declare the type of relationship. It links A to B, but it doesn’t say whether A belongs to B, whether A is an example of B, whether A extends B or whether A is a prerequisite for B. For a human user browsing, context is usually enough to infer the relationship. For an AI agent that processes the content graph to extract answers, the ambiguity has a direct cost in the quality of the result.
Semantic information architecture resolves that limit by declaring the type of relationship in the content’s metadata, independently of what the visible text says. Schema.org has been offering a standard vocabulary to do exactly that in the web context since 2011: the Course, EducationalOrganization and CollegeOrUniversity types make it possible to declare, in the page’s code, what kind of entity is being described and what attributes define it. An agent that reads that metadata doesn’t need to infer whether a page is about a program or a faculty: it knows with certainty, because it’s declared.
In June 2026, Google Cloud published the Open Knowledge Format (OKF), an open, platform-agnostic specification that takes this same idea one step further: instead of declaring knowledge in the markup of individual web pages, it proposes packaging it as a set of interconnected files —each describing an entity, its attributes and its relationships with other entities— that any agent can read directly, without the mediation of a crawler and without depending on the content being published at a public URL.
Machine Experience: the third layer of the audience
For most of the web’s history, the design of an institution’s digital presence answered to two audiences: the human users who browsed directly, and the crawlers of the search engines. UX served the former; technical SEO served the latter.
Contemporary information architecture adds a third layer of audience, which at Griddo we call Machine Experience: the AI agents that access an institution’s knowledge to answer queries, make decisions or execute tasks on behalf of a user. This audience doesn’t read web pages. It consumes structured knowledge. And its expectations about that structure are radically different from those of a human user.
A human user tolerates ambiguity because they have cultural context, inferential capacity and judgment to evaluate what they read. An AI agent works exactly within the limits of what is declared. What isn’t structured isn’t available. And what is available but poorly classified produces, at best, imprecise answers —and at worst, misinformation the system presents with the same apparent confidence it uses for everything else.
The 18% of candidates who have already ruled out a university based on what AI told them made that decision on the basis of the knowledge the institution did, or did not, make available to the systems that mediate that decision. It’s not a marketing problem or a reputation problem in the traditional sense. It’s an architecture problem.
The chain: from ontology to Machine Experience
The complete sequence has four links.
The first is the ontology: the decision about which entities exist in the university domain and what relationships are possible between them. This decision precedes any tool and any platform. It’s a conceptual decision about how the institution understands its own knowledge.
The second is the taxonomies: the classification systems that operationalize the ontology, assigning each entity a coherent position within the whole. A well-built taxonomy allows the institution’s various systems —the CRM, the web portal, the LMS, the admissions system— to talk about the same program with the same vocabulary, without each system describing it differently.
The third is content architecture: the structure of pages, records, relationships between documents and semantic metadata that materializes the ontology and taxonomy decisions into the institution’s digital presence. This is where schema.org, structured types and formats like OKF become concrete tools.
The fourth is Machine Experience: the result of having correctly built the three previous links. An institution that has precisely declared what its programs are, how they’re classified and what relationships they have with one another produces knowledge that AI agents can consume with confidence. That institution appears in the answers. The one that hasn’t done it, doesn’t.
Why this is a strategic decision, not a technical one
The usual objection to this kind of argument is that it’s a problem for the IT team or the web team, not for marketing leadership or academic leadership. It’s an understandable objection, and it’s wrong.
A university’s ontology —the definition of what a program is, how it’s distinguished from its variants, what relationships it has with other entities in the institution— cannot be built by the technical team alone. It requires a decision about the business model: does the university understand its graduate degrees as differentiated products or as variants of a unified catalog? Do programs belong to faculties or to cross-cutting fields of knowledge? Is employment data an attribute of the program or of the graduate’s profile? These questions have no technical answer. They have a strategic answer, and the answer the institution gives determines how its architecture is built and, therefore, what it can do with AI.
The universities that understand this before the rest will build an advantage that is structurally hard to copy in the short term, because knowledge architecture isn’t installed in a weekend. It’s designed, built, maintained and improved iteratively. The starting point matters, and the time to begin is before the pressure of becoming invisible turns it into an emergency.



