This is the second Skill for the Network-aware Daily arXiv Research Briefing Agent. It consumes paper metadata from Skill A and produces research-network analysis for daily briefings.
- Builds a coauthorship graph from paper authors.
- Builds a paper-similarity graph from embeddings or a TF-IDF fallback.
- Runs Louvain and Label Propagation community detection.
- Computes PageRank, betweenness, and degree centrality.
- Identifies bridge authors with
betweenness * (1 - clustering_coefficient). - Exports JSON plus static PNG visualizations.
python -m pip install -e ".[dev]"The core implementation uses only NetworkX and Matplotlib. PyVis is optional:
python -m pip install -e ".[interactive]"skill-b-network analyze --input examples/papers.json --output outputs/network_analysis.jsonLocal module form:
python -m skill_b_network.cli analyze --input examples/papers.json --output outputs/network_analysis.jsonThe command writes:
outputs/network_analysis.jsonoutputs/artifacts/coauthorship_network.pngoutputs/artifacts/paper_similarity_network.pngoutputs/artifacts/community_method_comparison.pngoutputs/artifacts/interactive_network.htmlif PyVis is installed
Required top-level fields:
run_id: run identifier used in reports.papers_metadata[]: paper records.
Required paper fields:
paper_idarxiv_idtitleabstractauthorscategoriespublished_at
Optional fields:
embeddings:{paper_id: [float, ...]}.history_snapshots: previous-day snapshots for future dynamic metrics.graph_params: overrides for defaults.
Default graph params:
{
"k_neighbors": 5,
"sim_threshold": 0.55,
"coauthor_weight": "count",
"community_methods": ["louvain", "label_propagation"],
"random_seed": 42,
"top_n": 10,
"visualization": {
"max_nodes": 60,
"label_top_n": 30,
"depth": 1,
"seed_nodes": 8,
"min_component_size": 1
}
}Visualization params affect the PNG/HTML artifacts only; the analysis JSON still contains the full graph.
max_nodes: maximum nodes drawn in one static network image.label_top_n: highest-ranked nodes that receive text labels.depth: neighbor expansion depth from the highest-ranked seed nodes when the full graph is too large.seed_nodes: number of central seed nodes used before depth expansion.min_component_size: hides tiny disconnected components in images when set above1.
The output JSON contains:
graphs: graph statistics for coauthorship and paper-similarity graphs.coauthor_edges: weighted author-author edges.paper_sim_edges: weighted paper-paper similarity edges.communities: method-specific assignments, community counts, and modularity.centrality: PageRank, betweenness, and degree centrality values.top_influencers: top authors and papers by PageRank.top_bridges: top bridge authors with explainability fields.emerging_communities: paper community summaries for the daily briefing.network_stats: repeated graph stats for downstream compatibility.artifacts: generated figure paths.warnings: recoverable data-quality and fallback messages.
python -m pytestSkill A should pass its ranked paper metadata and embeddings to this Skill. If Skill A has not implemented embeddings yet, this Skill still runs by using TF-IDF cosine similarity over title + abstract.
Skill B does not fetch arXiv data and does not summarize papers. It only analyzes metadata and vectors provided by the upstream retrieval/summarization layer.
- Author disambiguation is simple normalized-name matching.
- Louvain and Label Propagation are internal community estimates, not ground truth.
- Dynamic metrics require historical snapshots; without them,
growth_7d,novelty, andbridge_pressurearenull.