I do not have time nor knowledge to do a full-fledged "what is API" CEP.
Based on this discussion, I will gate the visibility of system_metrics behind a system property for now, by default it will not be visible. I believe this is the minimum we can do with maximum effect of not exposing this when it is obviously unstable / the subject of further improvements. It does not seem to me we are confident enough to say that system_metrics virtual keyspace is stable. On Thu, Sep 10, 2026 at 2:20 AM Yifan Cai <[email protected]> wrote: > > FYI, the CEP-65 thread > (https://lists.apache.org/thread/s8dsl4j1n95vpqtgxxzs20phs2g39m5j) is > discussing the same API lifecycle question for the shared utils library > instead of vtables. > > Same tension showing up in two places suggests the community wants a clear > API contract in general, not just for metrics. Might be worth one > project-wide policy instead of deciding it twice. > > - Yifan > > On Wed, Sep 9, 2026 at 12:47 PM Maxim Muzafarov <[email protected]> wrote: >> >> As an example from the OpenSearch ecosystem, plugins marked with the >> @Experimental annotation must be explicitly enabled in the YAML >> configuration and are disabled by default. This also makes their use >> quite visible from an operations perspective, since the configuration >> is easy for devops to grep when they need to understand the current >> cluster state or investigate a problem. >> >> ENABLE DEBUGGING / ENABLE METRICS is also a reasonable approach. >> However, it may be slightly less visible operationally because it is a >> runtime operation rather than something recorded in the static >> configuration (even if it's stored in TCM). >> >> Example from my recent experience: >> >> @Experimental >> public class IpfixSource implements Source<Record<Event>> >> >> experimental: >> enabled_plugins: >> source: >> - IpfixSource >> >> On Wed, 9 Sept 2026 at 13:18, Bernardo Botella >> <[email protected]> wrote: >> > >> > I would like to echo what David mentions here of the exposed vtables being >> > exposed to be considered part of the public APIs, therefore the >> > expectation should be that anyone can start building on top of them with >> > some guarantees of them not being broken. >> > >> > I also agree that, part of handling those expectations, can be to have >> > some of those APIs marked as experimental/prone to change/use them at your >> > own risk. For that, I don’t know if relying only on adding a EXPERIMENTAL >> > word to the comments field should be the way to go. Maybe we need to just >> > take a step back and consider them in the same boat we consider any other >> > API, which should have similar expectations and guarantees that are >> > consistent across the project. >> > >> > Forgetting for a moment that these APIs are vtables, I would think that >> > for all APIs across the project we need to be able to specify the >> > contract, and whether they are stable, experimental, or any other >> > terminology that help those using them making informed decisions. Also, >> > like with any other APIs, we need a process to be able to deprecate them. >> > >> > I have been trying to look for something project wide, and I don’t think >> > we have an actual policy around APIs. The closest I could find is this (1) >> > discussion from 2024 around experimental flagging. I think it would be >> > good to retake this conversation and end up with some good policies for >> > APIs the project can follow. >> > >> > Now, coming back to the EXPERIMENTAL flag for the vtables, if that’s what >> > we want to do “per policy” for an API, then we can discuss the best way to >> > do so? >> > >> > https://lists.apache.org/thread/9ptpkd3yymy5wok137c5jytysw373v52 >> > >> > >> > From: Štefan Miklošovič <[email protected]> >> > Date: Wednesday, 9 September 2026 at 07:27 >> > To: [email protected] <[email protected]> >> > Subject: Re: [DISCUSS] Future of system_metrics virtual keyspace >> > >> > > If they can see it then it should be assumed fair game; so maybe block >> > > them from seeing it until they agree to a set of shared rules? >> > >> > Yes, this is how I see it too. I think we are in a position that we >> > might declare all current vtables as API-stable - of course minus >> > Accord ones as Benedict mentioned, and we can make metrics >> > experimental as well. >> > >> > I scanned how vtables were evolving and we never removed anything, it >> > really is "stable" in that regard, what we have ever done was that we >> > were only adding new columns, never removing them. Maybe in one case >> > we changed the type of a column but otherwise it was addition only. >> > >> > So, now we declare that Accord + Metrics are experimental and hide >> > them until they are not. >> > >> > How the declaration of this should look like: >> > >> > 1) adding a comment into their CQL schemas that this is experimental >> > and probably subject of change >> > 2) a user would need to explicitly enable them via system properties, >> > your ENABLE METRICS is not a bad idea per se but I think that it is >> > just a "syntactic sugar" and a system property is just enough at this >> > point. >> > 3) we document this, in NEWS.txt, that this set of tables are >> > experimental and hidden. I think that is quite fair and enough, it is >> > expected that a user is reading this documentation, or at least >> > should. >> > >> > Once they are not in an experimental state anymore, we promote them to >> > be stable. It would need to be further clarified what that actually >> > means - what are the deprecation rules around this or if we have to >> > support it forever. >> > >> > On Tue, Sep 8, 2026 at 7:37 PM David Capwell <[email protected]> wrote: >> > > >> > > > My view is that virtual tables should not automatically be treated as >> > > > API-stable. They are for debugging / operator interaction, and intend >> > > > to expose internal implementation-specific state that is liable to >> > > > change across minors. >> > > >> > > If you expose them via JMX then we fail the build if you break the API >> > > as we need a stable API for operators… vtables are exposed in a much >> > > more “public” space so im not sure why the opposite should hold true. >> > > >> > > It also doesn’t make sense to me. If we can break the API in a patch >> > > release and ignore our deprecation process, then operators can’t use the >> > > APIs… then who are we building them for? In the example that started >> > > this thread there was desire to stop scraping JMX and use a vtable, but >> > > if we can break vtables when we feel like it then it would be dangerous >> > > to depend on the table so we logically should stick to JMX as its >> > > stable… then why do we have the vtable to begin with (I am not saying to >> > > drop the table, im just using it as a example against the argument)? >> > > >> > > We have similar issues with CMS/TCM tables… there was a desire to ask >> > > clients to move to the TCM peers table as its less queries on startup >> > > and more up-to-date… but if vtables are unstable we need to rely on the >> > > old ways. For tablets work we need a new vtable to expose the data >> > > placement, but again if vtables are unstable then its illogical for >> > > clients to touch it which becomes a blocker for tablets work... >> > > >> > > > We therefore must either avoid exposing internal state and >> > > > significantly hamper their utility, or else we must reject API >> > > > compatibility. >> > > >> > > This has been brought up before and I do agree with having an ability to >> > > denote that a table isn’t a stable API, and I would strongly agree with >> > > a solution to allow this. But this must be clear to a user else it’s an >> > > impossible situation for everyone. >> > > >> > > If you look at other projects you have flags you can call to get access >> > > to internal state and experimental features, you can create a CEP for >> > > this as it would be a way to have experimental and unstable apis >> > > >> > > cqlsh> ENABLE DEBUGGING; >> > > cqlsh> select * from system_unstable.why_is_repair_acting_up; >> > > >> > > You could even use this to allow specific feature >> > > >> > > cqlsh> ENABLE METRICS; >> > > cqlsh> select * from system_unstable.metrics. >> > > >> > > If we hide tables by default and expose a way to opt-in, we can define >> > > rules around them. We could have tables in experimental and be free to >> > > change them every major until we harden the API, in which case we >> > > promote it to the top level API. We could also use this to create >> > > tables we never intent to make stable; stuff that leaks internal state >> > > so it changes with that internal state. >> > > >> > > > My preference would be to standardise on a project policy of >> > > > defaulting virtual tables to API unstable unless explicitly declared >> > > > as stable (for programmatic access) >> > > >> > > And how would users ever discover this? They use CQL to find tables, >> > > they see it has the data they need, they build automation using it… then >> > > we break them when they upgrade and loose their trust. >> > > >> > > From a user’s point of view, why should they care about the >> > > implementation details of a table? If it’s disk backed vs in-memory why >> > > should they care? If they can see it then it should be assumed fair >> > > game; so maybe block them from seeing it until they agree to a set of >> > > shared rules? >> > > >> > > >> > > > On Sep 8, 2026, at 6:41 AM, Štefan Miklošovič <[email protected]> >> > > > wrote: >> > > > >> > > > I agree and also do not think that vtables should be automatically >> > > > treated as API-stable, should be "case by case" as yours are. >> > > > >> > > > But some vtables seem to be queried / parsed already and it will cause >> > > > a breakage as commentators on CASSANDRA-21539 reported, not sure how >> > > > to go about it, if we should codify which vtables are considered >> > > > API-stable and which are experimental, maybe just by putting its >> > > > experimental status into CQL table description (into "comment") or by >> > > > gating it behind a system property or similar. >> > > > >> > > > If metrics vtables are to be changed and we do not want to cause >> > > > confusion we might hide them by default and turn it off as Accord has >> > > > it. >> > > > >> > > > On Tue, Sep 8, 2026 at 3:03 PM Benedict Elliott Smith >> > > > <[email protected]> wrote: >> > > >> >> > > >> My view is that virtual tables should not automatically be treated as >> > > >> API-stable. They are for debugging / operator interaction, and intend >> > > >> to expose internal implementation-specific state that is liable to >> > > >> change across minors. We therefore must either avoid exposing >> > > >> internal state and significantly hamper their utility, or else we >> > > >> must reject API compatibility. >> > > >> >> > > >> To avoid an earlier argument about this, Accord virtual tables are >> > > >> simply off by default, so that the user must read the commentary that >> > > >> they are not API stable when enabling them. It may be suboptimal for >> > > >> users to realise this mid-incident though, but I cannot promise API >> > > >> compatibility for deep internal state that may cease to exist >> > > >> entirely. >> > > >> >> > > >> My preference would be to standardise on a project policy of >> > > >> defaulting virtual tables to API unstable unless explicitly declared >> > > >> as stable (for programmatic access), and - if it makes some people >> > > >> happy - to report a client warning on first access to such a table. >> > > >> >> > > >> >> > > >> >> > > >> On 2026/09/08 12:51:01 Maxim Muzafarov wrote: >> > > >>> Hi Stefan, >> > > >>> >> > > >>> Thank you for bringing this topic up. I think there is still some >> > > >>> room >> > > >>> for improvement here. >> > > >>> >> > > >>> Out of the options you outlined, I don't think the first option >> > > >>> excludes the second one. My preference would be to improve the API in >> > > >>> the 6.0 release, while still keeping it experimental for at least one >> > > >>> major release. This would give us some time to collect feedback from >> > > >>> real-world usage before treating the API as stable. >> > > >>> >> > > >>> We already had a discussion about experimental virtual tables and the >> > > >>> rules around them: >> > > >>> >> > > >>> [DISCUSS] Adding experimental vtables and rules around them >> > > >>> https://lists.apache.org/thread/xlv5rodt9v77rzrqssp3p63yjg0b88v4 >> > > >>> >> > > >>> A few additional thoughts from my side: >> > > >>> >> > > >>> 1. >> > > >>> If we want to change the UX or the schema of the metrics virtual >> > > >>> tables, I think it is better to do the larger changes in 6.0. This >> > > >>> way, users could expect only smaller or mostly cosmetic changes in >> > > >>> later releases rather than a complete redesign. >> > > >>> >> > > >>> At the same time, I think the expected usage patterns need to be >> > > >>> defined more clearly. Saying that there are "usability issues" does >> > > >>> not give us much direction unless we describe how users are expected >> > > >>> to query, filter, and export metrics. Otherwise, some of the proposed >> > > >>> changes may become a matter of preference rather than solving a >> > > >>> concrete problem. >> > > >>> >> > > >>> 2. >> > > >>> There is also a pending improvement waiting for a reviewer, which did >> > > >>> not receive the attention it deserved: >> > > >>> https://issues.apache.org/jira/browse/CASSANDRA-19666 >> > > >>> >> > > >>> This change improves the efficiency of bulk metrics exports >> > > >>> (benchmarks are attached to the issue). The issue also describes in >> > > >>> more detail the problems with keeping all metrics in a single large >> > > >>> collection, especially when querying or exporting them in bulk. >> > > >>> >> > > >>> I would therefore prefer to address the larger known UX and schema >> > > >>> issues in 6.0, while still keeping the API in an experimental state >> > > >>> for some time afterwards. This would give us more flexibility to make >> > > >>> smaller adjustments based on actual usage, rather than committing too >> > > >>> early to the current schema. >> > > >>> >> > > >>> On Mon, 7 Sept 2026 at 12:24, Štefan Miklošovič >> > > >>> <[email protected]> wrote: >> > > >>>> >> > > >>>> There is quite a thread on dev Slack discussing the implementation >> > > >>>> of >> > > >>>> virtual tables for metrics. Recently, there was also a lot of >> > > >>>> discussion on (2) where we initially removed a column in a virtual >> > > >>>> table which received a lot of pushback saying that it is a public >> > > >>>> API >> > > >>>> and we can not remove a column like that. >> > > >>>> >> > > >>>> Okay, fair enough. >> > > >>>> >> > > >>>> But when thinking about what was written in Slack about metrics and >> > > >>>> how vtables are modeled, I think that it is currently pending to be >> > > >>>> re-modelled quite a lot. We will very likely need to do this if we >> > > >>>> want to support e.g. more UX-friendly querying and scrapping of >> > > >>>> these >> > > >>>> metrics. Various people (Caleb, David) seem to be not satisfied with >> > > >>>> how system_metrics look like. (Please correct me if I am wrong). >> > > >>>> >> > > >>>> However, when 6.0 is out and we do not do anything about this, then >> > > >>>> by >> > > >>>> what was said in 21539 means that we will not be able to restructure >> > > >>>> vtable metrics and we will need to have yet another set of metric >> > > >>>> vtables on top of what we ship? >> > > >>>> >> > > >>>> What happens with virtual metrics tables after 6.0 when we will need >> > > >>>> to (massively) re-work their schemas? >> > > >>>> >> > > >>>> Possible options: >> > > >>>> >> > > >>>> 1) focus on its restructuralization before 6.0 is out because after >> > > >>>> that it will be effectively set in stone as that is our public API >> > > >>>> 2) mark it as experimental and object of further re-modelling >> > > >>>> 3) keep it as it is and never change it >> > > >>>> 4) ??? >> > > >>>> >> > > >>>> Regards >> > > >>>> >> > > >>>> (1) https://the-asf.slack.com/archives/CK23JSY2K/p1716324924466269 >> > > >>>> (2) https://issues.apache.org/jira/browse/CASSANDRA-21539 >> > > >>> >> > >
