Skip to main content

PageRank

SQL function: cugraph_pagerank

Official cuGraph reference: C API

Rank vertices by the stationary probability of a damped random walk that follows outgoing edges.

Signature

cugraph_pagerank(table_name [, src_col, dst_col [, weight_col [, options_json]]])

Relation inputs

The first positional argument names a registered edge table or view (the edges role). Parenthesized relation subqueries are not accepted; metadata validation uses the same registered name.

Vertex ID types

The edges relation declares the accepted vertex-ID domains. Numeric calls preserve the existing numeric schema. When logical string support is declared, Utf8, LargeUtf8, and Utf8View endpoint columns share one logical domain; their vertex-identity outputs are canonicalized to Utf8.

DomainAccepted endpoint inputsOutput contract
Numeric edge endpointsInt32, Int64The numeric output schema is used for numeric calls.
Logical string edge endpointsUtf8, LargeUtf8, Utf8ViewVertex identity columns are canonicalized to Utf8; scores, distances, counts, coordinates, and opaque labels remain numeric.

The native mapping type is Int64. Call-specific output schemas come from gpu_validate_call.

Logical string side-input limitations:

  • edge ID columns and edge-ID predicate side inputs are not supported for logical string graphs

Scalar arguments & JSON options

Positional scalar arguments

src_col and dst_col name the edge endpoint columns; both are optional and default to src and dst.

ArgumentTypeRequiredDefaultNotes
weight_colUtf8|nullnooptional edge weight column for graph construction when supported by the algorithm; semantic effect: edge weights affect algorithm results when provided

JSON options

OptionTypeDefaultConstraintsDescription
alphaFloat640.85min 0; max 1
epsilonFloat640.00001> 0
max_iterationsUInt32100min 1

Graph construction options

Graph construction follows the shared defaults (directed=true, renumbering, python_cugraph policy) documented in Graph Construction Options.

Output schema

ColumnTypeNullableDescription
vertexInt64|Utf8noVertex receiving the PageRank score.
valueFloat64noPageRank score for the vertex.

These are generic descriptor schemas; validate the call to get the concrete, table-specific output schema.

Examples

These examples run on the citation network demo dataset (4.9M papers, 45.6M src-cites-dst edges).

PageRank versus raw citation counts

PageRank over the full graph, joined back to paper metadata. Importance flows through citations: a paper cited by foundational papers outranks a paper with more, but shallower, citations.

SELECT p.title, p.year, p.n_citation, r.value AS pagerank
FROM cugraph_pagerank('citation_edges', 'src', 'dst') r
JOIN papers p ON p.paper_id = r.vertex
ORDER BY r.value DESC
LIMIT 5;
titleyearn_citationpagerank
Finite automata and their decision problems19591,4010.00139
The Mathematical Theory of Communication194948,3270.00134
The reduction of two-way automata to one-way automata19592240.00126
The complexity of theorem-proving procedures19714,5920.00073
A mathematical theory of communication194822,1220.00066

The first and third papers have modest raw counts (1,401 and 224 citations), but the papers citing them are themselves foundational results, and PageRank propagates that structure. The full call — 45.6M edges, GPU graph build, 4.1M scores, join, and sort — returns in about 1.5 s.

Window functions over the result

The output of a cugraph_* function is a plain relation, so ROW_NUMBER() applies directly to it. This query finds where the paper that introduced PageRank ranks, by its own algorithm, among 4.9 million papers.

WITH ranked AS (
SELECT vertex, value, ROW_NUMBER() OVER (ORDER BY value DESC) AS rank
FROM cugraph_pagerank('citation_edges', 'src', 'dst'))
SELECT r.rank, p.title, p.year, r.value
FROM ranked r JOIN papers p ON p.paper_id = r.vertex
WHERE r.vertex = 2066636486;
ranktitleyearvalue
115The anatomy of a large-scale hypertextual Web search engine19980.000143

Rank #115 of 4,894,081.

SQL defines which graph the GPU sees

The first argument is any relation name, including a view. Joining the edge list to papers on both endpoints restricts the graph to citations within a chosen era, and PageRank then ranks the most important papers within that window — here, the pre-2000 literature.

CREATE VIEW edges_pre2000 AS
SELECT e.src, e.dst
FROM citation_edges e
JOIN papers ps ON ps.paper_id = e.src
JOIN papers pd ON pd.paper_id = e.dst
WHERE ps.year BETWEEN 1901 AND 2000 AND pd.year BETWEEN 1901 AND 2000;

SELECT p.year, p.title
FROM cugraph_pagerank('edges_pre2000', 'src', 'dst') r
JOIN papers p ON p.paper_id = r.vertex
ORDER BY r.value DESC
LIMIT 6;
yeartitle
1959Finite automata and their decision problems
1959The reduction of two-way automata to one-way automata
1949The Mathematical Theory of Communication
1974The Design and Analysis of Computer Algorithms
1958Preliminary report: international algebraic language
1963Machine perception of three-dimensional solids

Automata theory, information theory, the classic algorithms textbook, and the ALGOL report: the view's WHERE clause restricts the graph to one era, and the algorithm re-ranks the field within it.

Limitations & lifecycle

No algorithm-specific limitations.

Validate before running

Dry-run validation checks registered relation metadata, column presence, static dtypes, and options only; it does not scan edge data, construct a graph, or prove source-vertex existence:

SELECT * FROM gpu_validate_call(
'cugraph_pagerank',
'{"schema_version":1,"relations":{"edges":{"table":"target_edges"}},"options":{"src_col":"src","dst_col":"dst"}}'
);

See GPU Function Catalog API for the full gpu_validate_call contract.