These functions measure how well some proposed structure describes a network. Unlike the intrinsic properties in measure_features, each takes a structure from the user — a core-periphery mark, or a partition of the nodes — and returns how closely the observed network corresponds to it:
net_by_core() measures the correlation between a network
and a core-periphery model with the same dimensions.
net_by_factions() measures the correlation between a network
and a component model with the same dimensions.
net_by_modularity() measures the modularity of a network
based on nodes' membership in defined clusters.
net_by_linkdensity() measures the partition density of a network
based on ties' membership in defined clusters.
net_by_divergence() measures how far a network is from an ideal
network, such as a core-periphery, complete, or star network.
net_by_inconsistency() measures how far a partition's blocks depart from
ideal block types.
These are the natural companions to the node_in_*() functions, which
propose a structure; these say how good that proposal is.
Where a partition is expected but none is given, the network is
partitioned into two using node_in_partition().
Note that they are not on a common scale, and do not all run in the same direction, so they are not interchangeable:
| measure | compares the network against | range | better |
net_by_core() | a core-periphery model | -1 to 1 | higher |
net_by_factions() | a components model | -1 to 1 | higher |
net_by_modularity() | the partition's communities | -0.5 to 1 (at the default resolution) | higher |
net_by_linkdensity() | the partition's communities of ties | -1/3 to 1 | higher |
net_by_divergence() | an ideal network | 0 to 1 | lower |
net_by_inconsistency() | ideal block types | 0 upwards | lower |
Compare partitions using one measure at a time.
net_by_core(
.data,
mark = NULL,
variant = c("correlation", "ident", "ndiff", "diff"),
coreness = NULL,
direction = c("all", "out", "in"),
method = NULL
)
net_by_factions(.data, membership = NULL)
net_by_modularity(.data, membership = NULL, resolution = 1)
net_by_linkdensity(.data, membership = NULL)
net_by_divergence(
.data,
ideal = manynet::create_core,
variant = c("hamming", "jaccard", "portrait"),
mark = NULL,
membership = NULL
)
net_by_inconsistency(.data, membership = NULL, blocks = c("nul", "com"))A network object of class stocnet, igraph, tbl_graph, network, or similar.
Internally any of these will be coerced to an efficient implementation.
For more information on possible coercions, see e.g. manynet::as_stocnet().
A logical vector indicating which nodes belong to the core.
Character string naming which variant of the measure to compute, where more than one definition of the same quantity is in use. The variant chosen is reported when the result is printed.
Which method to use to calculate nodes' coreness. One of "correlation", "rich", "transition", or "hub"; see method_coreness for what each does. By default NULL, which uses "rich" for a weighted, directed, or two-mode network, since it is the only method that reads those properties directly, and "correlation" otherwise.
One of "all" (the default), "out", or "in". For a directed network, "out" scores nodes on the ties they send and "in" on the ties they receive. Ignored for undirected and two-mode networks.
Deprecated. The former spelling of variant.
Still accepted, but warns; please use variant instead.
A character string naming an existing node attribute in
the network, or a categorical vector of the same length as the number of
nodes in the network where each element indicates the group membership of
the corresponding node.
While this may often be a vector created using node_in_*() functions,
it can be any character vector that assigns nodes to groups or categories.
A proportion indicating the resolution scale. By default 1, which returns the original definition of modularity. The higher this parameter, the more smaller communities will be privileged. The lower this parameter, the fewer larger communities are likely to be found.
The ideal network to compare against. Either a network object,
or a function that builds one from the network, such as any of the
manynet::create_*() or manynet::generate_*() functions.
A function is called with the network as its first argument, and also
given mark or membership where it takes them.
By default manynet::create_core().
For other arguments, pass an anonymous function,
e.g. \(x) manynet::create_lattice(x, width = 4).
A character vector of permitted ideal block types,
or a list-matrix giving the permitted types for each block position.
By default c("nul", "com"), which is structural blockmodelling.
See the section below.
A network_measure numeric score.
The object also carries the measure it computed, the range its values
can fall within, and whether and how those values were normalized.
These are shown as a one-line header when the object is printed.
Where a measure offers a choice between several ways of counting the
same thing, it also carries the variant it used.
All can be retrieved with attr().
net_by_modularity() counts a tie however it is signed, as a census does,
so where the network is signed each tie is read by its magnitude.
net_by_core() and net_by_factions() fit the network to an ideal built
by a manynet::create_*() function, and those build one layer at a time.
A multilevel network holds two, so there is no single ideal to fit it to
and both stop rather than compare unlike shapes.
Take one layer first, e.g. with manynet::to_mode1() or
manynet::to_uniplex().
For net_by_core(), which of the following to use to calculate the fit of
the core assignment to a core-periphery model.
"correlation" calculates the correlation between the empirical network and
an ideal typical network, and "ident" calculates the Euclidean distances
between the same.
"ndiff", however, calculates how distinct the core and periphery groups are
based on the difference in coreness scores between the least core-like
member of the core and the most core-like member of the periphery.
"diff" is similar to "ndiff", but multiplies the raw "ndiff" score by the
square root of the size of the core, thus penalising large cores.
net_by_core() calculates the Pearson correlation between the given
network, where the nodes in the core are assigned by some given mark, and
an ideal typical core-periphery network with the same number of nodes in
the core and the periphery.
Where mark is not given, it is calculated with node_is_core(), to
which the coreness and direction arguments are passed. For a directed
network the fit itself is measured on the symmetrised network, since the
ideal it is compared against is symmetric.
Modularity measures the difference between the number of ties within each community
from the number of ties expected within each community in a random graph
with the same degrees. At the default resolution it ranges between
-0.5 and +1; a higher resolution can push it further below that floor.
Modularity scores approaching +1 mean that ties only appear within
communities, while negative scores mean that ties appear between
communities more often than chance would predict.
A score of 0 would mean that ties are half within and half between communities,
as one would expect in a random graph.
Modularity faces a difficult problem known as the resolution limit
(Fortunato and Barthélemy 2007).
This problem appears when optimising modularity,
particularly with large networks or depending on the degree of interconnectedness,
can miss small clusters that 'hide' inside larger clusters.
In the extreme case, this can be where they are only connected
to the rest of the network through a single tie.
To help manage this problem, a resolution parameter is added.
Please see the argument definition for more details.
net_by_linkdensity() measures the partition density of a membership of
the network's ties, such as that from tie_in_community(),
which is used where no membership is given.
Each community's density is the number of its ties beyond those that a
tree over its nodes needs, as a share of the ties that would make those
nodes a clique.
The partition density is the mean of these, weighted by the number of
ties in each community (Ahn et al. 2010).
It is 1 where every community is a clique, 0 where every community is a
tree, and negative only where the ties of a community do not connect.
Ties between the same two nodes count once, and ties without a
membership are left out.
net_by_divergence() measures how far a network is from some ideal
network, from 0 where the network is the ideal to 1.
There are three variants:
"hamming" (the default) is the share of possible ties on which the
network and the ideal disagree. This is the graph edit distance, counting
only tie additions and deletions, normalised by the number of possible ties.
A missing tie that is also missing in the ideal counts as agreement.
"jaccard" is one minus the share of ties in either network that are in
both. Unlike "hamming", it ignores ties missing from both, so a sparse
network does not look close to a sparse ideal just because both are sparse.
For the same reason it is always 1 against an empty ideal,
so use "hamming" there.
"portrait" is the portrait divergence of Bagrow and Bollt (2019): the
Jensen-Shannon divergence between the two networks' distributions of how
many nodes each node reaches at each distance.
It compares structure at all scales and needs no correspondence between
the two networks' nodes, so the ideal may even be of a different size.
"hamming" and "jaccard" compare the two networks tie by tie, and so
need the ideal's nodes to correspond to the network's.
This holds for an ideal network of the same dimensions and node names,
and for manynet::create_core(), manynet::create_components(),
manynet::create_filled(), and manynet::create_empty(), which are built
from the network's own mark or membership, or do not depend on the
order of nodes.
It does not hold for, e.g., manynet::create_star() or
manynet::create_ring(), where which node is the hub or comes next is
arbitrary, nor for the random manynet::generate_*() functions.
There the portrait divergence is used instead, and the result reports
the variant used. Note that an ideal from a manynet::generate_*()
function is random, so it returns a different value on each call.
Where a directed network is compared tie by tie with
manynet::create_core() or manynet::create_components(), which build
one direction only, both are first symmetrised, as in net_by_core().
net_by_divergence() and net_by_inconsistency() both return how far a
network is from an ideal, and both run lower-is-better,
but they differ in what that ideal is:
net_by_divergence() compares against one ideal network, and returns
a value between 0 and 1.
net_by_inconsistency() compares against the best of a set of ideals.
Each block may take whichever of the permitted types fits it best,
so the data choose the image. Some of those types, such as reg, are
satisfied by many different blocks rather than by one, so there is no
single network to compare against. Its value runs from 0 upwards.
Where every block of a membership is best fitted as complete on the
diagonal and null off it, net_by_inconsistency(blocks = c("nul", "com"))
equals the "hamming" divergence from
manynet::create_components(membership = membership).
In short, use net_by_divergence() where you can name the ideal network,
and net_by_inconsistency() where you can only name the permitted
block types.
A blockmodel proposes that a partition reduces a network to a small number
of positions, so that every block — the ties running from one position to
another — is of some simple ideal type.
net_by_inconsistency() measures how far the network departs from that proposal,
by counting the ties that would have to be added or removed to make every
block ideal, normalized by the number of cells.
Lower is better: 0 means the partition fits perfectly.
This is a distance from an ideal image rather than a measure of fit — hence the name, and hence its running the opposite way to the rest of this page. Three consequences are worth knowing:
Its complement is not a proportion. The criterion mixes units: nul
and com count cells, while reg counts empty rows and columns, and all
are divided by the cell count. So do not read \(1 - x\) as the share of
the network that the blockmodel gets right.
It is not bounded above by 1. That holds only for cell-counting
vocabularies such as c("nul", "com"). With reg permitted it can exceed
1 — on ison_adolescents, blocks = "reg" over singleton positions
reaches about 1.57.
The vocabularies behave very differently at fine partitions. Giving
every node its own position scores 0 under c("nul", "com"), since each
block is then a single cell and trivially ideal, but scores its worst
under "reg", since each block then has an empty row and column.
For a correlation-scaled, higher-is-better reading of the common structural
case, see net_by_factions(). The two are related but not equivalent:
net_by_factions() fixes the image — complete on the diagonal, null off it
— whereas net_by_inconsistency(blocks = c("nul", "com")) lets each block take
whichever of the two ideals fits it better, and so is more permissive.
For how it relates to net_by_divergence(), which also runs
lower-is-better, see the section on divergence and inconsistency.
The ideal types are:
nula null block, containing no ties.
coma complete block, containing every possible tie.
rega regular block, in which every row and every column has at least one tie, though not necessarily all of them.
rdo, cdoa row- or column-dominant block, containing at least one complete row or column.
dnc"do not care": a block left unconstrained.
blocks is a vocabulary rather than an assignment: each block is scored
at the lowest inconsistency of any permitted type, and the results summed.
Any subset may be given, and the two conventional choices are
c("nul", "com") for structural equivalence and c("nul", "reg") for
regular equivalence.
Note that permitting more types can only lower the criterion, since each block gains more ways to be satisfied. The size of the vocabulary is therefore itself a modelling choice, and criterion values are comparable across partitions only when the same vocabulary is used for each.
For fully generalized blockmodelling, pass a g by g list-matrix
naming the types permitted at each position separately,
e.g. reg on the diagonal and nul off it for a "cohesive positions"
model.
Borgatti, Stephen P., and Martin G. Everett. 2000. “Models of Core/Periphery Structures.” Social Networks 21(4):375–95. doi:10.1016/S0378-8733(99)00019-2
Newman, Mark E.J. 2006. "Modularity and community structure in networks", Proceedings of the National Academy of Sciences 103(23): 8577-8696. doi:10.1073/pnas.0601602103
Murata, Tsuyoshi. 2010. "Modularity for Bipartite Networks". In: Memon, N., Xu, J., Hicks, D., Chen, H. (eds) Data Mining for Social Network Data. Annals of Information Systems, Vol 12. Springer, Boston, MA. doi:10.1007/978-1-4419-6287-4_7
Ahn, Yong-Yeol, James P. Bagrow, and Sune Lehmann. 2010. "Link communities reveal multiscale complexity in networks". Nature 466(7307): 761-764. doi:10.1038/nature09182
Bagrow, James P., and Erik M. Bollt. 2019. "An information-theoretic, all-scales approach to comparing networks". Applied Network Science 4: 45. doi:10.1007/s41109-019-0156-x
Doreian, Patrick, Vladimir Batagelj, and Anuska Ferligoj. 2005. Generalized Blockmodeling. Cambridge: Cambridge University Press. doi:10.1017/CBO9780511584176
Other features:
measure_features
Other measures:
measure_assort_net,
measure_assort_node,
measure_breadth,
measure_broker_node,
measure_broker_tie,
measure_brokerage,
measure_central_between,
measure_central_close,
measure_central_degree,
measure_central_eigen,
measure_central_tie_between,
measure_central_tie_close,
measure_central_tie_degree,
measure_central_tie_eigen,
measure_closure,
measure_closure_node,
measure_cohesion,
measure_core,
measure_diffusion_infection,
measure_diffusion_net,
measure_diffusion_node,
measure_diverse_net,
measure_diverse_node,
measure_features,
measure_fragmentation,
measure_hierarchy,
measure_periods
net_by_core(ison_adolescents)
#> # Core-periphery correlation [-1, 1]
#> [1] -0.133
net_by_core(ison_southern_women)
#> # Core-periphery correlation [-1, 1]
#> [1] -0.274
net_by_factions(ison_southern_women)
#> # Factional correlation [-1, 1]
#> [1] 0.485
net_by_modularity(ison_adolescents,
node_in_partition(ison_adolescents))
#> # Modularity [-0.5, 1]
#> [1] 0.155
net_by_modularity(ison_southern_women,
node_in_partition(ison_southern_women))
#> # Modularity [-0.5, 1]
#> [1] 0.31
net_by_linkdensity(ison_adolescents,
tie_in_community(ison_adolescents))
#> # Link density [-0.3333333, 1]
#> [1] 0.35
net_by_divergence(ison_adolescents, manynet::create_filled)
#> # Hamming divergence, normalized [0, 1]
#> [1] 0.643
net_by_divergence(ison_adolescents, manynet::create_filled,
variant = "jaccard")
#> # Jaccard divergence, normalized [0, 1]
#> [1] 0.643
net_by_divergence(ison_adolescents, manynet::create_star)
#> # Portrait divergence, normalized [0, 1]
#> [1] 0.821
net_by_inconsistency(ison_hightech, node_in_regular(ison_hightech))
#> # Blockmodel inconsistency [0, Inf)
#> [1] 0.221
# a regular-equivalence vocabulary instead of a structural one
net_by_inconsistency(ison_hightech, node_in_structural(ison_hightech),
blocks = c("nul", "reg"))
#> # Blockmodel inconsistency [0, Inf)
#> [1] 0.0762