Resource Data Federation: let's discuss

We should make some time to talk about the challenges of resource data federation – by which I mean, enabling multiple directory maintainers to collaborate on data management across different information systems.

I think there are technical questions in here (like how to enable cooperation among people with different answers to questions like “what is a service?”) and also governance questions (like how to divide up responsibilities, incentives, conflict resolution etc) although the former might blur into the latter.

We certainly will need more than one conversation to really make progress, but maybe we could just start with a fishbowl discussion among people who have some experience. I’m happy to facilitate.

@klambacher has agreed to participate. I’m sure @skyleryoung will be interested. I also know of at least a couple of groups of people interested in pursuing this path who might not be here, but who I’d invite to listen in and bring questions.

NB: I want to wait for @MikeThacker-iStandUK to get back before picking a date because he’s asked for this for a while.

Who else is interested? What should be on our agenda? What questions would you want such a conversation to consider?

I’m definitely in. When is Mike back?

When is Mike back?

I’m back now.

OK great. I assume that since we’re spanning pacific coast to uk time, we’re looking at the 11a ET hour – sound right?

@klambacher @MikeThacker-iStandUK @skyleryoung and anyone else who might be interested, let me know if you have any days in which 11a ET normally doesn’t work.

Next week i’ll schedule something after reaching out to the folks in the field who asked me for this convo.

OK I have folks from at least one state eager to join this conversation, so i’m ready to start scheduling.

I’m looking at 11a Eastern during any Monday, Tuesday, Thursday, or Friday during the last week of March or first week of April. I’ve created a poll here: Doodle

Please fill it out! And let me know if there’s anyone you think i should specifically invite. thank you :slight_smile:

OK we have picked March 25th at 11am Eastern for a discussion about resource directory data federation.

We’ll invite some opening remarks from @klambacher and also @skyleryoung. Then we’ll have discussion.

Please help set the agenda in advance: what questions come to mind when you consider the challenges and opportunities for different organizations using different information systems to collaborate in management of shared resource directory data? Please send in the questions you would most like to discuss, and I will try to shape an agenda that is worthwhile to everyone :slight_smile:

Hi folks – thanks so much to @klambacher for sharing with us the story of her epic journey through resource data federation. I’ve got notes here (feel free to add or clarify).

I know that several attendees had questions first and foremost about politics and trust and organizational strategy issues, so it was great to hear Kate’s lessons learned about those things. I don’t think we got very far into the technical challenges as pertains to some of our questions about evolving the specification – so I assume we may want at least one more round of discussion on this.

I will first check to see whether the same day/time will work for most folks next week or the week after – so, April 1st or 8th at 11a Eastern. Please let me know if that won’t work for you.

And please also share your observations and questions for the next agenda!

Good Afternoon, I am interested in participating in the next conversation(s) if possible please.
To your question about dates, I am available on the dates you mentioned as option (April 1st or 8th at 11a Eastern).
Thanks
Erika (No Wrong Door Virginia)

Hi folks – we’re scheduled to resume this conversation at 11a Eastern on tuesday next week.

Now that we’ve heard a real-world story of a federation experience from Kate, I’d like this next conversation to open up a bit to explore the technical and tactical challenges.

I know we have technical questions about structuring identifiers among distributed sources, and provenance metadata.

Also there’s the question of how to establish “official” or “core” or “golden” records that are reliably verified by a designated steward in a way that still enables others to add their own custom data.

We also have questions about structuring governance, and incentives – i.e. how organizations can be equitably compensated for distributed contributions.

What are your questions? Please let me know so we can do a bit of work in advance to shape the agenda :slight_smile:

@bloom it would be good to hear:

  • what motivates publishers to keep up-to-date with the latest version of HSDS/HSDA?
  • what additional properties do people add to help their federation? (e.g. Kate mentioned tags to decide what appeared in what outputs)

Hi folks – Thanks to those who joined for the time yesterday.

Notes are here: Standing Technical Committee Meeting Notes – Open Referral - Google Docs

I went through to bold the parts that I thought were most significant. I would welcome more comments and key takeaways from you all.

I know that this question pops out at me, and seems relevant to the feedback loop design question, and in turn the metadata specifications:
can/should there be nested (vertical) levels of responsibility? So vertical rather than horizontal redundancy – i.e. local/topical responsibility for record stewardship, overseen/bottom-lined by higher-level umbrella stewardship. (Is vertical vs horizontal the right frame?)

Should we pick up there? Are there other questions you want to discuss next?

I’m going to send out a calendar item to reconvene and continue the discussion for two weeks from now. Let me know if you’d like to keep this going and if that would work for you.

OK notes from this conversation below and here, please feel free to help clarify / add / organize. This is a bit circular but a good discussion. we have a couple of specific questions we could pick up next. what do you want to discuss next?

Incentives –

“Local stewards are more likely to have relationships and are able to get people on the phone at the organization. More experienced data managers, however, understand the needs for the funders and how to craft a highly specialized and high quality record.”

Kate: “reconciling these two strengths is both necessary and possible.” experienced data managers can pass down their knowledge / build up data quality, while locals can build and maintain relationships.

“Duplicate record management is not always bad, in-fact some overlap in management of a singular resource drives more rich information and more frequent updates. Reconciling those records to match (suggest things are the same) and then map (human editing and confirmation of the matching)…” Creating a golden record set that distills this (?)

Who should get paid for what?

  • Locals / domain experts/stewards for managing records on a per-record basis.
  • Regional/state-level for bottom-lining comprehensive data, quality assurance, etc.
  • Producing “opinionated” data sets for specific
  • Participation in cooperative processes i.e. governance activities
  • CBOs should get paid for providing services! This is important; it may or may not be relevant.

First two bullets can be gamed / played unequally. Organizations (local OR high-level/utility) might be incentivized to maximize their benefits in ways that aren’t in the best interests of users or the network as a whole. Example of incentivizing the wrong things: paying more for complex records leads to needless complexity.

^ but these challenges can be addressed through governance.Contract based payments with incredibly specific restrictions around what records are being compensated has really been proven to be effective in disincentivizing cheating in the system. Agreements on fair terms, and monitoring and conflict resolution processes to manage and evolve the agreements over time. These can also yield technical safeguards, though that has to follow agreements. “Tight standards and inclusion policies, especially coupled with auditing, make it so that each steward can keep their own records, that’s their business, and whether they get paid or not is on the basis of standards.”

Get clear on what we are trying to incentivize?

  • Reliability.
  • Responsiveness
  • Relationships.

Are there different thresholds – some kinds of data are objective, can be collected with relatively minor levels of expertise; some kinds are specialized and need expertise.

Skyler: “Paying a per record fee when we are at a very early stage can be tricky. Knowing the actual cost of production of a record and having that be itemized is critical.”

Kate: “We need to find a way to recognize the value of checking a record even when significant information does not churn. We need to investigate measuring how an entire record is checked.”

Sasha: Time is tricky. a lot of the work will not be reflected in the metadata captured in digital tools. “Sometimes you need to chase an organization for 2 months, occasionally an organization will be really on top of it and hand the details to you proactively” “Phone conversations are currently offline and aren’t always able to be traditionally programmatically tracked. Beyond that, if you get sucked into a 50 minute conversation about a random side tangent with one organization” (David side thought: it helps build rapport and may help them be more helpful in the future so this is not a bad thing)

Kate: re expecting people to meet standards to participate, there is a level of expertise that involves personnel investment and specialization that can be difficult where the role is part-time or volunteer or not a primary responsibility.

Chris: “Being able to find the money in the system to support even just one FTE at a local level can really be differential for the quality of the data in that locality. Being able to earmark that funding and then allowing the local organization to get creative with how they spend that to drive results, can be a useful blend”

Can there be thresholds for compensation,

  • Thresholds for simple vs complex
  • Thresholds for minimal compliance with standards vs full compliance

We do want to track costs. But is it worth tracking all the metadata in order to compensate on a granular level?

Skyler:

​​1. Incentives are connected to Responsibilities, and

  1. We don’t assume that data driven metrics are the exclusive method for establishing compensation

Also, I’m personally thinking in terms of ratios of payment, not necessarily cost of record maintenance. I’m inclined to decouple those concepts for at least.

Kate –

Assuming we get funding for data maintenance:

A funder especially gov don’t want to pay a bunch of people. They want to fund one. So it naturally falls to a model where they fund a top-level organization, and the top-level picks the next level, and those pick their local partners. And funding flows differently across levels.

What are non-monetary incentives? Often at the lowest levels, the exchange might need to be other than money: training, software, access to data + analytics, support resources… what else?

Need roles for Auditing / review / standards-setting. This makes it hard to pass on top-level funding to lowest level, cuz it eats a lot up.

Who should pay for what?

There might be different kinds of payers in the market –

Funders that just want to pay for production of data (government?)

Funders that want to pay for specialized products – curation

We want proper staff… and we also want those local relationships that may be more informal.

So – Utility concept of bottom-line / overarching responsibility to aggregate and publish

Steward concept for designated responsibility that adheres to standards.

Community partner concept for leveraging local relationships.

A utility plays stewardship roles as well as administration, auditing, and curation.

A community partner may play a stewardship role, if they have the capacity / interest in assuming a higher level of responsibilities.

Should the concept of a ‘curator’ be distinct from (even if co-assigned with) ‘steward’? What about ‘auditor’?

Can / should we try to unbundle stewardship responsibilities for complex organizations, i.e. to distribute responsibilities across the service / location level?

David: “One thing that I think will be significant to investigate for capturing the value of the administration, education, and licensing is the travel agent and travel agency business model, and how those organizations interact with the major airlines, hotels, cruise lines, etc. Its a random spot of knowledge I have a decent amount of detail on, and the structure is really relevant to what I have learned”

Thanks again to those who joined. I went through the notes and tried to pull out a set of takeaways, itemized at the top of the notes and pasted below.

Next week many of us will be at the Inform USA conference, in which these converastions can continue informally.

Is there interest in re-starting the conversation in June? Say June 3rd or 10th at 11aEastern? If so, what should the agenda be?

see below and chime in – thanks!

Takeaways:

Roles:

  • Utility: responsible for aggregate, publish, and quality control. May have specific front-line stewardship responsibilities in addition to bottom-line / umbrella responsibilities; may also play role of curator and auditor for other stewards.
  • Steward: responsible for a specified level of data management of a given set of resources.
  • Auditor: responsible for assessing the quality of data managed by a given set of stewards, to ensure compliance with standards. (performed by Utility
  • Curator: responsible for ensuring consistency in subjective elements of resource records, especially category/taxonomy

There should be designated stewardship responsibilities for the ‘core’ part of the record. That said, taxonomies might benefit from being managed (curated) centrally.

For aggregated vs unbundled service information – i.e. Programs that might involve multiple subsidiary services – Perhaps bundled program record is shared broadly, and discrete service records are kept locally and/or provided for a fee.

^ Question: Can modern transformer tools help accommodate a both/and balance between loose and strict? So that systems can have it both ways with tooling to automate bundling/unbundling.

Style guide : data standard :: style template : exchange profile

in order to get alignment on complex orgs, develop a ‘template’ for certain kinds of organizations (like gov agencies) to specify how information about them should be structure.

In order to get alignment on specific data point to share for specific purposes/users, develop an exchange profile – “Export profile” or “import profile” – that specifies particular fields for particular purposes, and sets up the crosswalking to be automatable.

Who should get paid for what?

  • Locals / domain experts / stewards should get paid for managing records –
    • on a per-record basis?
  • Regional/state-level utility should get paid for aggregating comprehensive data, quality assurance, etc.
  • Producing “opinionated” data sets for specific consumers
  • Participation in cooperative processes i.e. governance activities
  • Some local work (non-standard contributions, unstructured input) can be contributed by community partners who might not get paid but can benefit in nonmonetary ways (training, support)
  • CBOs should get paid for providing services! This is important overall; but it may or may not be relevant to data management strategy.

Hi folks – we had good conversations last year about Federation (as per the notes above), and we heard in the Technical Committee meeting today that it may be time again to pick the thread back up.

I think an agenda for this could probably go in various directions, so I want to hear: who is interested in exploring the question of federation further, and what questions would you want to bring to that conversation?

Once we’ve heard from a few people, I will find us a time :slight_smile:

In response to @bloom and following a commitment I made at the last technical group meeting, here’s a list of topics that I’d like to see covered in documentation about combining federated datasets:

  • minimum data compliance expectations (which parts of HSDS)
  • data quality expectations/agreements
  • processes for addressing data quality problems
  • common identifiers for organizations and other elements in data from different publishers
  • de-duplicating records from different sources

I’m sure there are many other points raised by Kate. If someone can point me to notes from the sessions we had with here, I’ll review them and add what I consider important in the UK. I expect our Greater Manchester work to involve aggregating data from multiple sources.

OK we’ve locked in Tuesday next week at 11am Eastern for our conversation about federation. IF you have not received a calendar invite but you WOULD like to attend, please reply here or to me and I can add you.

As a reminder, here was the context that prompted the revival of this series. And here are the notes from the last of the series of these discussions last year (previous discussion notes are in subsequent subject headers in that doc starting with “Breakout”)

As for the agenda, the context that led to this discussion was around the differences between “data guides” (in which one party is specifying the contents of data published for many to consume) and “export profiles” (in which one party is structuring data that is specifically designed to be consumed by a designated other party or set of parties) and “import protocols” (which is a name that we haven’t agreed upon, it’s just my attempt to describe what I’m hearing has to happen by the receiving party).

Our overall goal is to work toward shared understanding of the most important considerations that have to be addressed in order for federation to actually work – and one specific objective is identifying which if any elements of this that might be A) specified as part of the standards, B) supported with specific tooling, or C) just promoted through documentation.

So @MikeThacker-iStandUK some of your bullet points above relate to this, but I think our interest is in grounding technical patterns within organizational context.

Would welcome any other suggestions for agenda items. Review the notes and let us know what you want to discuss!

@klambacher any examples or reference materials or even just outlines that might help us “see” your experience with export profiles etc would be helpful. Let me know if some time beforehand to chat it through might help.

I won’t be able to join the call on Federation, because I look after my daughter on Tuesdays. But here are my initial thoughts, based on my experiences with ActivityPub (a standard for Federated social web activities e.g. Mastodon, Pixelfed, Lemmy etc.)

Make the federated protocol a separate standard to HSDS

Make the federated protocol a seperate standard. HSDS is a great model for Service data, but it risks being overloaded if we pile a bunch of models/protocols for Federation in there.

My ideal approach would have three layers:

  1. HSDS models: this provides the “what”, basically service data.
  2. Models describing federated activity: “create this”, “update this”, “delete this”, “redact that action” etc. These models directly reference the HSDS schemas
  3. A protocol for describing how to exchange layer 2: which endpoints, what HTTP request to perform, when to perform them.

In ActivityPub, the ActivityPub protocol leans on a separate “ActivityStreams” standard. ActivityStreams provides the models/vocabulary for objects such as an Activity (Create, Delete, Update, Undo, etc.). The ActivityVocabulary spec defines the types of object e.g. Article.

So for Open Referral, I’d want to keep layer 1 as HSDS. Layer 2 would provide another set of models which define which actions a server should perform, and Layer 3 would be the exchange protocol for exchanging them.

If Layer 2 was defined serparately then the specification for layer 3 can focus on its work of defining compliance. It will need to dictate how actors on the network discover each other, how these objects are exchanged and how actors should behave when they create/recieve each item.

Focus on server-to-server interactions

ActivityPub explicitly denotes two sets of protocol interactions: client-to-server (an app or interface speaking with a server), and server-to-server (different servers exchanging information and keeping up to date).

I would think that the server-to-server is the more important for service data federation. A client-to-server protocol allows people to “bring their own interface”, whereas I imagine interfaces for managing service data already exist and will be tightly coupled to whatever systems people are already using; the challenge is getting these systems to exchange information properly.

Race conditions and malicious behaviour

Federation, or distributed systems, can result in race conditions.

For example, someone using Server A might create a Service object (or update one) and then immediately delete it (or redact the update). Depending on the behaviour/speed of the network, there is a chance that the object which says “Delete this service” reaches Server B before the instruction to create the service object.

This means Server B may do nothing when told to delete something that isn’t there, and then receive the instruction to create a new service, which it then complies with naïvely. This is a real problem in ActivityPub world.

A lesson learned would be that the protocol should cover this and declare that compliant implementations MUST behave in a particular way when receiving a request to delete something that they haven’t created yet. Off the top of my head, this could be something like “store the request to delete and then periodically check that it can be acted upon, and then perform the action when you can” or “when you receive the request to create something, check that you don’t already have a request to delete it before creating it”.

There is always the potential for malicious behaviour on the federation, where a server sits on the federation and appears to be behaving nicely but never actually deletes any records from its database. This might be harmless in the context of open data, but since there exists a potential for sensitive data being leaked; it might cause issues. Even if the majority of the federation behaves appropriately; if someone accidentally leaks a sensitive address and then redacts it, a malicious server would be able to harvest that and process it store it in some way.

While Server A can refuse to send updates to Server B, if Server C hasn’t blocked them then Server B might be able to get Server A’s updates via Server C. Therefore, there needs to be tools/protocols in place for instances to block hostile actors, and prevent data percolating through the federation.

On a separate note, I would also make sure that authentication and security is covered in the protocol; how can we ensure that a request from Server A to delete some service record actually comes from Server A? ActivityPub handles this via HTTP Signatures, which are some cryptography that is performed by the server.

Inboxes and Outboxes are useful concepts

Activity Pub has concepts of Inboxes, Outboxes and Actors. I think server-to-server interactions for exchanging service data would mostly benefit from Inboxes and Outboxes.

Inboxes and Outboxes are basically just URLs which have specific behaviours associated with them. Inboxes allow servers to POST things (activities) to each other, and Outboxes allow retrieval via GET.

Actors might be less useful to the federated exchange of service data. ActivityPub has them because implementations are often concerned with managing groups of users and topics, each of which might have their own Inboxes and Outboxes. For service data, it might be sufficient that each server simply has its own inbox and outbox.

Implementation Agnostic

I’m sure I don’t need to say this, but writing anyway for completeness; the design of the protocol should be agnostic to implementation tech beyond what is required for exchanging information (HTTP and JSON). It shouldn’t matter to the protocol whether a server’s database is denormalised or normalised, or is built on the latest shiny cloud infrastructure vs held together by PHP and bubblegum as long as it can be proven to be compliant.

Testing compliance

Testing compliance will be more difficult than mechanically reading an openapi.json file and testing that HTTP GET works across the endpoints listed like we do for validating HSDS API feeds. There will need to be a custom test suite which tests behaviour of servers responding to requests for specific actions. I’m not sure how the internal audit of these systems would work, but a test suite could run a bunch of tests and then give its confidence on how well a server/instance/implementation is complying.

Thanks for all this, very interesting, although I suspect most of the conversation next week will fall above the level of exchange protocol (more like orgware)…

Meanwhile – could we just anticipate using ActivityPub?

Thanks for all this, very interesting, although I suspect most of the conversation next week will fall above the level of exchange protocol (more like orgware)…

Ah yeah I thought my reckons might be a bit technical-heavy; without direct experience of the challenges of federated service data at the organisation level I am definitely more suited to the “here’s how the protocol and data models might work” level of conversation.

Meanwhile – could we just anticipate using ActivityPub?

I think this is a reasonable first port of call. ActivityPub itself might not be directly suitable because it is explicitly designed for federated social web (Twitter-likes, Reddit-likes, Facebook/YouTube clones etc). However the ActivityStreams Core vocabulary (Create, Delete, Update, Undo) is a solid base to start with and might take us 80% of the way there. What would be left for us is to explicitly define which HSDS models should be getting exchanged (all of them, most likely) and how servers are expected to discover each other, endpoints, and behave when receiving/sending messages.