Cybernetics of AI Companies

原始链接: https://ai-cybernetics.grok.me

Hacker News new | past | comments | ask | show | jobs | submit login Cybernetics of AI Companies ( grok.me ) 3 points by measurablefunc 48 minutes ago | hide | past | favorite | discuss help Consider applying for YC's Winter 2027 batch! Applications are open till November 2. Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact Search:
相关文章

原文

Abstract

An AI company is modeled as an open, resource-dependent system whose observations, learning procedures, deployment actions, and governance rules close several feedback loops. A contextual -topos supplies a language for states, admissibility certificates, and coherent identification of implementations; categories and lenses retain the direction of irreversible operations. We give elementary results on certified update preservation, descent, the limits of observational regulation, and the insufficiency of semantic equivalence for identifying learning dynamics. At the differentiable level, an optimizer-dependent neural tangent kernel is derived exactly along gradient flow. Parameter jets, a neural tangent hierarchy, and an input-jet tangent kernel describe distinct extensions. Their order is independent of homotopy truncation. Every extensional optimizer on smooth objectives factors through their global holonomic jet representation, with runtime and oracle state retained; factorization through a jet at one point has a separate locality condition. A formal panoptic hypothesis specifies trace coverage and inequalities between observation, intervention, and governance capabilities. The construction combines identity types, univalence, and jets into a research program for the cybernetics of AI firms, without treating economic concentration, universal surveillance, or a fixed-kernel description of every neural system as mathematical consequences.

Keywords. Organizational cybernetics; homotopy type theory; -topos; lenses; neural tangent kernel; jets; constrained updates.

The company-level application is a proposed synthesis. Its starting hypotheses are that employee and customer traces can both become productive training data, that feedback changes the behavior being measured, and that data integration can create incentives for consolidation. These are conditional empirical and institutional claims. The mathematical results below concern explicitly specified models; they do not establish that every actual AI company collects all available traces or that a single company must emerge.

Wiener’s control-and-communication perspective [35], Ashby’s requisite variety [1], and Beer’s organizational cybernetics [2] motivate the separation of operation, coordination, adaptation, and policy. Modern categorical cybernetics supplies compositional interfaces and controllers [4, 27]; categorical learning supplies parametrized maps, learners, and reverse differentiation [9, 7]. Higher groupoids add a precise account of witnessed equivalence and its coherence. They are not substitutes for economic evidence or for differential equations.

Two further strands connect architecture and institutions: categorical deep learning studies algebraic architecture constraints [13, 12], while compositional game theory models interacting decision-makers [14]. Hedges explicitly proposes categorical-cybernetic analysis of AI-mediated markets and supply chains [16]. That proposal is a research agenda, not an established empirical law of AI companies.

Central question. Which observations and interventions can a firm perform, which specifications must those interventions preserve, and how do learning, governance, and the environment respond to one another? A model is useful when it makes these questions executable or falsifiable, rather than merely relabeling a company as a category.

2.1 Foundational conventions and notation

A category has objects, arrows with specified source and target, identity arrows, and associative composition. A functor preserves these data. A natural transformation consists of arrows between two functors’ values which commute with every source-category arrow. “Small” means that the objects and arrows belong to the chosen set-sized universe. reverses arrows, and denotes functors and their natural transformations. A groupoid is a category whose arrows are invertible. An endomorphism has equal source and target; an automorphism is an invertible endomorphism. A terminal object has one arrow from each object, and an initial object has one arrow to each object, with the corresponding mapping spaces contractible in higher categories. Products and limits are defined by their universal mapping properties.

For a precise higher model, a simplicial set is a functor

In the internal type language, a universe classifies the chosen small types. A dependent family

A presheaf of spaces is a functor

An open system exchanges inputs and outputs across a chosen boundary. A controller chooses inputs from observations and possibly memory; feedback feeds outputs into subsequent inputs. Governance specifies who may change controllers, objectives, interfaces, and admissibility rules. Resources are stocks or capacities consumed or replenished by operations. These terms acquire explicit maps, state variables, and authorization relations in the constructions below.

2.2 The identity tower

Work in intensional dependent type theory, with a chosen univalent universe when univalence is used. For

The groupoid model demonstrates that uniqueness of identity proofs is not a general consequence of intensional type theory [17]. Martin-Löf’s foundational presentation [26] and Voevodsky’s homotopy -calculus notes [34] provide foundational background. Grothendieck’s homotopical perspective [15] connects suitable models of higher groupoids to homotopy types.

Definition 2.1

Truncation levels

A type is a mere proposition, or -type, when any two inhabitants are equal. It is a set, or -type, when each of its identity types is a mere proposition. Inductively, it is an

Its universal property is

2.3 A topos of organizational contexts

Definition 2.2

Contextual firm semantics

Choose a small site . Objects of describe contexts such as a team, project, deployment, jurisdiction, or accessible interface; arrows describe restriction of context. The topology specifies which compatible local views jointly cover a context. Set

Firm states, interfaces, specifications, and equivalence witnesses are objects or structured objects internal to .

This is an actual -topos, obtained by accessible left-exact sheafification [25]. The selected site is a presentation of the semantics. Neither the bare organization nor an arbitrary category of neural networks automatically forms a topos. A -truncated part supports ordinary sheaf logic and the Mitchell–Bénabou language [22]; suitable universes and dependent type-theoretic structure support the higher identity language [32]. An arbitrary such model does not automatically supply every desired computational rule for a proof assistant.

A state in context is a generalized element

For differentiation, choose smooth parameter spaces and derivative operations. A smooth manifold is a Hausdorff, second-countable space covered by Euclidean charts with smooth transition maps; smooth maps are smooth in those charts. A concrete option is to replace the context site by a product with a site of Cartesian spaces and smooth maps, with the corresponding product topology, and to use smooth finite-dimensional charts for parameter objects. Probability is extra structure too. A standard Borel space is measurably isomorphic to the Borel space of a complete separable metric space. A Markov kernel

Remark 2.3

Three independent resolutions

Homotopy order , differential jet order , and access/capability resolution answer different questions. Here labels an observation-and-action interface, not a numerical dimension; it is unrelated to the cardinal used in the accessibility convention. A representable smooth sheaf is the functor

3.1 States and admissible transitions

Let the ambient state object be

where describes human roles and capabilities, data and provenance, models and parameters, optimizer/runtime memory, controllers, resources, and governance. Products may be replaced by dependent sums when, for example, an optimizer’s state depends on the architecture. Specify a predicate

Thus an inhabitant of contains both a proposed state and its validity witness. The predicate may include schema equations, authorization, resource bounds, deployment tests, and declared contractual conditions. Their concrete meaning must be supplied; a proof of the chosen predicate is not proof of every desirable social outcome.

Definition 3.1

Certified update

For an admissible input family

Proposition 3.2

Preservation under iteration

Suppose an initial state has a witness

Proof. The base case is

This formalizes constrained database updates and speculative changes. A valid rollback concerns the modeled state: deletion of already transmitted information, or reversal of an irreversible physical act, requires additional operations and cannot be assumed from this proposition.

For an ordinary categorical database schema , an instance is a functor

Proposition 3.3

Descent of certified operations

Let

Proof. Apply the sheaf equivalence between sections over and the limit of sections over the Čech nerve to the structured object of certified operations. This object is assembled from dependent sums, products, and the specified certificate family. The fiber over the given descent datum is contractible.

Independent local approvals are not enough. Agreement on shared data, interfaces, and higher coherence is part of the premise. For example, two resource requests may each fit a budget locally and exceed it jointly; they do not form a valid global resource certificate.

3.2 Interfaces, lenses, and feedback

A deterministic open module consists of a state object , an output map

More generally, has its own state and the environment contributes inputs; the full product state must then be retained. In a stochastic interpretation, updates are Markov kernels and composition uses their integral composition law.

For interfaces

The second component carries responses, requested changes, or cotangent signals back through the interface. For

This is the elementary bidirectional structure; dependent lenses permit response types to depend on the current output. Reverse differentiation sends a smooth map

Parametric maps and architecture constraints. A layer with parameters is a map

Algebraic architecture constraints can also be expressed through monads. In an ordinary category, a monad consists of an endofunctor and natural transformations

An algebra

The temporal operations in (2) are generally irreversible. They belong to a category of transitions or systems. Only equivalences belong to its maximal higher groupoid. In an -category higher morphisms are invertible; its ordinary -morphisms need not be. Reflexivity, symmetry, and transitivity of a relation alone do not provide the entire identity tower.

An executable trace

People and environmentobserveTraces and data

Traces and datatrainLearning and evaluation

Learning and evaluationreleaseDeployment and action

Deployment and actionintervenePeople and environment

Governance and resourcesauthorizeDeployment and action

Learning and evaluationmetricsGovernance and resources

Figure 1. A proposed firm-level feedback architecture. Each arrow has a typed interface and an explicitly chosen operational interpretation.

Univalence asserts that the canonical map

is an equivalence [32]. It permits transport of type-dependent properties and structures along a supplied equivalence. The identity structure includes automorphisms; univalence does not reduce all equivalence witnesses to one proof. Nor is it derivable just from a verbal appeal to Leibniz’s principle. The structuralist interpretation motivates a chosen foundation; it does not establish a uniqueness theorem for all possible foundations.

For firms or learners, equivalence must be defined on structured systems. An appropriate object includes

Example 4.1

Equal predictions, different learning speeds

Let

The models have the same realizable predictions and equivalent parameter spaces, but different time-dependent predictions from matched initial outputs.

Proposition 4.2

Kernel invariance requires optimizer geometry

Suppose

Proof. The chain rule gives

Reading figure · reparametrization

Same realizable map, two learning speeds, as in Example 4.1. Here

This gives a concrete requirement for a higher groupoid of learner presentations: equivalences of optimizer-equipped systems must preserve the relevant geometry. Gauge symmetries such as hidden-unit permutations may be organized into an action groupoid, and, with descent, a quotient stack . Its stabilizers record presentation automorphisms. A quotient is not automatically a smooth manifold, and a kernel descends only when the required invariance holds. Here a group action is a map

Example 4.3

Truncation does not identify realizations

For every

The point is the one-point homotopy type. The Eilenberg–Mac Lane space , for

so -truncation fails to preserve this homotopy pullback. Sheafification and homotopy truncation must therefore play distinct roles.

5.1 Finite models and continuous-time optimization

The notation means derivatives through order exist and are continuous; means this for every finite . Write for the Jacobian, for the -linear derivative, for the gradient in the declared inner product, and for its Hessian. A dot denotes a time derivative. A real symmetric matrix is positive semidefinite, written

Fix a dataset and concatenate a differentiable network’s training outputs into

Here is the number of parameter coordinates, the number of stacked output coordinates, and

Theorem 5.1

Optimizer-dependent tangent dynamics

Along (4),

The matrix

Proof. Multiply (4) by to obtain (6). For any ,

Reading figure · residual bound

The squared-loss claim in Theorem 5.1. If

For

Jacot, Gabriel, and Hongler established the foundational NTK viewpoint and a constant-kernel infinite-width regime under their hypotheses [19]. Linearized wide-network results [23] and tensor-program calculations for broad architectural classes [36] extend its scope. They do not imply that an arbitrary finite transformer or other deployed network has a constant kernel. Feature learning can change and ; lazy training is a particular scaling regime [5] in which the parametrized map remains close to its linearization at initialization over the specified training interval.

5.2 Kernel evolution and the tangent hierarchy

For , sufficiently differentiable and no explicit time dependence of ,

The evolution already involves second parameter derivatives, and derivatives of the optimizer geometry. For the special scalar-output Euclidean flow with

The labels

This is the recursive structure of the neural tangent hierarchy (NTH) [18]. These higher tensors are not, in general, positive semidefinite two-input kernels or fully symmetric tensors. Freezing a finite level is a closure approximation; the identities themselves do not bound its error. Huang and Yau obtain controlled truncations under specified smoothness, width, initialization, and data assumptions. Their approximation theorem is not asserted here for every modern architecture or optimizer.

6.1 Germs, the infinite jet tower, and holonomic sections

Fix an open parameter domain

Definition 6.1

Smooth germ and pointwise jet

The germ

with bonding maps forgetting the highest-order derivatives.

An inverse limit here is a compatible sequence under these forgetful maps; the germ direct limit identifies representatives agreeing after restriction. In Taylor coordinates the natural map

Definition 6.2

Global holonomic jet representation

The prolongation of is its jet section over the entire domain,

The assignment

6.2 Jet algebra and its limitations

In a chosen parameter chart, the order- parameter jet of a map is its Taylor polynomial modulo terms of degree

At a fixed point it records derivatives through order , not the entire map. Under coordinate changes these records transform by the chain rule. They can be organized into jet bundles, whose fibers collect the jets at each base point, and compositional differential structures [3, 30, 21]. Here

Equivalently, with unnormalized derivatives,

Lemma 6.3

Differentiation lowers finite jet order

For

Proof. If two representatives differ by , their derivatives differ by a multiple of , so they have the same

The infinite jet, or formal power series, admits the coefficient shift without this finite-order loss. Truncated polynomial representatives can be differentiated after choosing a representative, but that is a chosen approximation to the germ’s derivative. Regular jets do not automatically include Laurent series with a pole at the base point. Shared differential formulas are not an equivalence of all polynomial, Laurent, smooth, and formal theories.

6.3 A universal jet representation of optimization algorithms

Definition 6.4

Stateful objective-based optimizer

Let be a runtime state space containing current parameters and any memory, query history, population of candidates, or random seed used by the algorithm. Let contain declared auxiliary data, constraints, precision, schedules, objective presentations, and oracle access rules. A deterministic optimizer step is a possibly partial map

Code inspection and representation-dependent costs are explicit auxiliary data when they affect the algorithm. A randomized program may equivalently use a deterministic step with its seed in . The oracle is the interface answering objective or derivative queries; a black-box algorithm knows the objective through those answers rather than through a closed formula.

Theorem 6.5

Universal global-jet factorization

Every optimizer of Definition 6.4 on smooth objectives has a unique optimizer on global holonomic jets,

For partial maps its domain is transported by

Proof. The inverse identities in Definition 6.2 give the displayed factorization. Every in the holonomic space equals

This is the precise sense in which every specified objective-based optimization algorithm is an instance of the global jet construction. It includes derivative-free methods because the objective’s values are its zeroth-order component. The theorem is a representation result: infinite derivatives can be available without each algorithm actually computing or using them. It also holds with global order- holonomic sections for any

Proposition 6.6

Compositional equivalence of optimizer interfaces

Fix auxiliary data. Let

Proof. Identity steps return ; associativity follows from substitution. Since

Composition permits search, differentiation, filtering, acceptance, resource checks, and release gates to use one typed objective representation while retaining their separate runtime states.

Proposition 6.7

Criterion for a pointwise-jet optimizer

Fix and a total germ-local rule

Proof. A factorization is constant on the fibers of the jet map. Conversely, assign to each realizable jet the value of any representative germ. The premise makes this assignment independent of representative, giving the unique descended rule. The same reasoning applies to a partial rule precisely when definedness also descends.

Here germ-local means dependence only on agreement on a neighborhood; pointwise-jet determined means dependence only on the derivative data at the base point. These are distinct conditions. For example, choose a smooth function supported in with

Orders, queries, and optimizer state. Gradient methods read first-order jets of at their query points. Newton methods read second-order jets. A line search selects a step length along a declared direction by querying candidate values; it can be expressed through zeroth-order evaluations of the global jet section. Finite differences use several such evaluations to approximate derivatives. Quasi-Newton methods update an approximate Hessian or inverse Hessian from successive gradient and displacement pairs, so their memory belongs to . Momentum and Adam combine first-order queries with retained memory. Population and randomized search methods retain several candidate points and query values at each. A trust-region method optimizes a local Taylor model under a bound

For a composite loss

For nonsmooth or discrete objectives, classical infinite derivative jets may not exist. The universal value interface remains: a function

6.4 Second and higher-order optimizers

For

Newton-type updates use this second jet. A damped positive-definite approximation can supply

For an actual finite parameter step , Taylor’s theorem gives

with

6.5 An explicitly defined input-jet tangent kernel

The phrase “jet tangent kernel” does not specify a single universal object. For a precise version, take continuous inputs

Definition 6.8

Order-$r$ input-jet tangent kernel

For a parameter preconditioner , define the block kernel

The observable coordinates in this definition are unnormalized derivatives; using Taylor-normalized coordinates rescales the corresponding blocks.

For independent of input, differentiating the ordinary output block kernel gives

Proposition 6.9

Jet-observable dynamics

For training on

Proof. Apply the chain rule and the Gram-matrix argument of Theorem 5.1 to in place of .

This is appropriate for derivative matching and Sobolev objectives; Sobolev training is an established example [8]. A sampled Sobolev objective, for example, is

For discrete tokens there are no ordinary input derivatives in the token index; input jets can instead be taken in continuous embedding variables when that is the intended model. Parameter derivatives of logits remain meaningful wherever the parametrization is differentiable.

6.6 What no fixed kernel can determine

For momentum, optimizer memory is essential: matched parameters can have different velocities. For Adam [20], with gradient

Here

Here is standard Brownian motion, the specified drift vector, the noise-amplitude matrix, and the matrix trace. The Hessian correction is absent from deterministic gradient flow. ReLU, the rectified linear unit

As Definition 6.1 shows, the infinite jet at one point need not determine a smooth germ. Theorem 6.5 instead uses the global holonomic section and the full optimizer state. Analyticity restores local determination by a convergent Taylor series; it does not remove optimizer or environmental dependence. The justified differential statement is the chain-rule identity for differentiable parametrized models under specified update laws. A single fixed “jet kernel” does not determine all training, inference, deployment, and organizational dynamics.

7.1 Capability-relative observation

For an actor , choose an observation map

Theorem 7.1

Observation-fiber obstruction to regulation

Let

Conversely, a supplied selection from these intersections defines a regulator. In constructive type theory, the converse requires an actual dependent selection, not merely an assertion that each intersection is nonempty.

Proof. If regulates, then

In particular, if indistinguishable states need incompatible actions, no extra computation applied to the same instantaneous observation can solve the problem. Memory or new sensors can change the observation object. If a finite disturbance set requires a different unique action for each disturbance, any regulating observation must distinguish every disturbance:

7.2 Employee traces, customer traces, and the panoptic hypothesis

An observational trace is a finite sequence of typed interaction events with declared times, participants, and outputs. An employee trace comes from work performed in an employee role; a customer trace comes from use of a product or service in a customer role. One person can occupy both roles. Data provenance specifies origin, transformations, and the permissions attached to the data. These traces become inputs to a learner only through a specified collection and data-processing map. This observational use of “trace” differs from an executable operation sequence.

Let be existing data, a newly captured work trace, and a future prediction target, all modeled as finite random variables on a declared probability space. The Shannon conditional entropy, in bits, is

Log loss for a predicted distribution is

This follows directly from the definition of conditional mutual information and nonnegativity of conditional relative entropy. It gives a precise condition under which internal traces have predictive value. It does not establish equal value to customer data, a monetary valuation, or an incentive to capture every trace: acquisition cost, provenance, representativeness, permissions, and objective choice remain independent variables.

For worker and customer trace families

Definition 7.2

Observation refinement and authority

For a common state domain , observation maps

The channels and rights concern a declared organizational interface, rather than every fact an actor knows or every action an actor can perform. Controller inspection and challenges to updates are explicit observation or governance commands when included in this interface.

Definition 7.3

Trace coverage

Choose a finite operational state abstraction , including the relevant firm, environment, and interaction histories. Let

The number

Recoverability is equivalent to

Definition 7.4

Panoptic architecture and panoptic hypothesis

Let contain the state domain, roles, observation maps, trace observables, weights, actions, and governance relations just defined. It is a panoptic architecture at threshold , written

  • for every
  • for some

For a declared collection of firms and a fixed assessment period, the panoptic hypothesis is the empirical assertion

This is an operational definition inspired by the asymmetry of observation and authority in the historical account of panopticism [10]. It specifies the mathematical meaning of the term in this note. The threshold, role weights, trace scope, recovery errors, and authorization evidence are part of the hypothesis and must be reported. Positive predictive value alone proves none of its three conditions. A model can have full trace coverage and fail the definition when all action and policy-change rights are equal. Conversely, authority concentration does not establish high trace coverage. Controllability means the ability to reach declared target states through allowed inputs; observational predictability is a separate property.

Proposition 7.5

Panoptic status is invariant under structured equivalence

In the finite model, bijective changes of state, observation, trace, action, and policy coordinates preserving roles, weights, maps, and authorization relations preserve

Proof. A recovery map transports by composing with the observation bijection’s inverse and the trace bijection. Hence exactly the same role instances contribute to

In higher contextual semantics a recovery witness has the internal type

Mere recoverability is its -truncation, while the full type retains maps and coherence witnesses. Numerical coverage above uses a finite, decidable abstraction; a general higher recovery proposition is not automatically a decidable Boolean. An observation-only equivalence which forgets permissions need not preserve panoptic status. The structured equivalence must retain the authority data as well.

Deployment also changes data: write

7.3 Composed decision-makers and shared providers

A company contains and connects decision-makers with distinct objectives. The set-based open-game formalism [14] represents a component

An equilibrium in context is a strategy with

Hedges’s research proposal [16] highlights the dependence created when apparently separate economic decision-makers use a shared upstream AI provider. A minimal probabilistic example makes the distinction precise. Suppose a finite latent provider regime affects two otherwise conditionally independent action channels. Then

under the stated conditional-independence assumptions. This mixture generally does not factor into the two marginal channels. For

8.1 Learning stability does not imply firm stability

An equilibrium is a state with zero time derivative. It is asymptotically stable when all sufficiently close initial states remain close and converge to it. For the linear systems below this is equivalent to every eigenvalue having negative real part. A kernel mode is an eigenvector

One possible frozen-kernel interpretation is

Proposition 8.1

Two-mode closed-loop criterion

The equilibrium of (16) is asymptotically stable exactly when

Proof. The characteristic polynomial is

Reading figure · two-mode loop

Proposition 8.1. The linearization

Increasing an isolated learning rate, or making its loss decrease, does not by itself control and . Delays, switching releases, and nonlinear saturation require a richer model. A categorical wiring diagram specifies composition; it does not provide the missing stability estimates.

Resources can be made explicit through, for example,

with an available energy stock and a financial balance, together with material stocks and capacity constraints. Power is energy per unit time. A viability region includes resource floors as well as technical and institutional conditions. Maintaining that region is a control problem with disturbances; parameter-loss minimization proves only the narrower statement in Theorem 5.1.

Beer’s viable-system roles suggest a decomposition into operational model teams, coordination of shared infrastructure, allocation and audit, adaptation to the environment, and policy/identity [2]. This is a modeling correspondence, not evidence that an AI company is a biological organism or already a viable system in Beer’s technical sense.

8.2 A conditional consolidation comparison

Suppose compatible data objects admit a pushout

This defines the pushout as the gluing of the span through

Let

Here the terms are avoided duplication and complementary-data gains, while the terms are the respective additional costs. Under this accounting, integration is preferred for this objective precisely when the right side is positive. This is a conditional comparison, not a prediction: independent objectives, diseconomies, limited resources, failure concentration, and governance can reverse the sign. No date or inevitability of industry-wide consolidation follows from categorical universal properties or from (15).

The framework supplies a starting point with several measurable extensions. For a first case study, choose one workflow from employee or customer interaction to training, evaluation, release, and subsequent observations. Record its state variables, time delays, optimizer memory, permissions, resource costs, and rejection paths. Specify before analyzing updates. Identify which equivalences are intended to preserve predictions only, and which must also preserve optimizer geometry, costs, or control rights.

Estimate empirical tangent-kernel modes or Jacobian-vector products on a declared sample; materializing a full kernel is often unnecessary. Track kernel change, jet approximation error over actual step sizes, and sensitivity of behavior to release and data-selection decisions. Estimate the coupled response terms in (16); test whether broader observations resolve the obstruction in Theorem 7.1. These measurements separate information capture, effective intervention, and governance.

Four theoretical extensions follow naturally:

  • Compositional certificates. Give module contracts and prove that their shared-resource and interface witnesses assemble into the descent datum of Proposition 3.3. Arbitrary local tests do not suffice.
  • A moduli object of learners. Define equivalences preserving specified forward behavior, losses, optimizer geometry, and permissions. Compute stabilizers and ask which kernels and costs descend to its quotient stack.
  • Multiobjective and stochastic governance. Replace a single scalar company objective by stakeholder-indexed objectives and a specified mechanism for selecting actions. Add Markov kernels, delays, and resource-aware policies.
  • Jet closure with error control. Derive architecture- and optimizer-specific truncation estimates, rather than importing a smooth fully-connected NTH bound into an unrelated runtime.

Relating cocycles to homotopies requires a specified complex or derived moduli problem. A cochain complex has vector spaces and maps

in degrees . Here a Lie group is a smooth manifold with smooth group operations, is its tangent space at the identity, and

Finally, a one-point model exists for many purely equational algebraic theories, including the theory of groups, but not for every consistent theory. The theory of nontrivial fields, with

  1. [1]

    W. Ross Ashby. An Introduction to Cybernetics. Chapman and Hall, London, 1956. Requisite variety: Chapter 11.

  2. [2]

    Stafford Beer. Brain of the Firm. John Wiley and Sons, second edition, 1981.

  3. [3]

    R. F. Blute, J. R. B. Cockett, and R. A. G. Seely. Cartesian differential categories. Theory and Applications of Categories, 22(23):622–672, 2009.

  4. [4]

    Matteo Capucci, Bruno Gavranović, Jules Hedges, and Eigil Fjeldgren Rischel. Towards foundations of categorical cybernetics. Electronic Proceedings in Theoretical Computer Science, 372:235–248, 2022.

  5. [5]

    Lénaïc Chizat, Edouard Oyallon, and Francis Bach. On lazy training in differentiable programming. Advances in Neural Information Processing Systems, volume 32, 2019.

  6. [6]

    Roger C. Conant and W. Ross Ashby. Every good regulator of a system must be a model of that system. International Journal of Systems Science, 1(2):89–97, 1970.

  7. [7]

    G. S. H. Cruttwell, Bruno Gavranović, Neil Ghani, Paul Wilson, and Fabio Zanasi. Categorical foundations of gradient-based learning. Programming Languages and Systems (ESOP 2022), LNCS 13240, pages 1–28. Springer, 2022.

  8. [8]

    Wojciech Marian Czarnecki, Simon Osindero, Max Jaderberg, Grzegorz Swirszcz, and Razvan Pascanu. Sobolev training for neural networks. Advances in Neural Information Processing Systems, volume 30, 2017.

  9. [9]

    Brendan Fong, David I. Spivak, and Rémy Tuyéras. Backprop as functor: A compositional perspective on supervised learning. 34th Annual ACM/IEEE Symposium on Logic in Computer Science, pages 1–13, 2019.

  10. [10]

    Michel Foucault. Discipline and Punish: The Birth of the Prison. Vintage Books, 1995. Translated by Alan Sheridan. “Panopticism,” pp. 195–228.

  11. [11]

    Tobias Fritz. A synthetic approach to Markov kernels, conditional independence and theorems on sufficient statistics. Advances in Mathematics, 370:107239, 2020.

  12. [12]

    Bruno Gavranović. Fundamental Components of Deep Learning: A Category-Theoretic Approach. PhD thesis, University of Strathclyde, 2024.

  13. [13]

    Bruno Gavranović, Paul Lessard, Andrew Joseph Dudzik, Tamara von Glehn, João Guilherme Madeira Araújo, and Petar Veličković. Position: Categorical deep learning is an algebraic theory of all architectures. Proceedings of the 41st International Conference on Machine Learning, PMLR 235, pages 15209–15241, 2024.

  14. [14]

    Neil Ghani, Jules Hedges, Viktor Winschel, and Philipp Zahn. Compositional game theory. Proceedings of the 33rd Annual ACM/IEEE Symposium on Logic in Computer Science, pages 472–481, 2018.

  15. [15]

    Alexander Grothendieck. Pursuing stacks. Circulated manuscript, 1983. Revised transcription edited by Mateo Carmona with Ulrik Buchholtz, 2021.

  16. [16]

    Jules Hedges. AI safety meets value chain integrity. Cybercat Institute, December 11, 2023. Research proposal.

  17. [17]

    Martin Hofmann and Thomas Streicher. The groupoid interpretation of type theory. In Twenty-Five Years of Constructive Type Theory, Oxford Logic Guides 36, pages 83–111. Oxford University Press, 1998.

  18. [18]

    Jiaoyang Huang and Horng-Tzer Yau. Dynamics of deep neural networks and neural tangent hierarchy. Proceedings of the 37th International Conference on Machine Learning, PMLR 119, pages 4542–4551, 2020.

  19. [19]

    Arthur Jacot, Franck Gabriel, and Clément Hongler. Neural tangent kernel: Convergence and generalization in neural networks. Advances in Neural Information Processing Systems, volume 31, pages 8571–8580, 2018.

  20. [20]

    Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. International Conference on Learning Representations, 2015.

  21. [21]

    Ivan Kolář, Jan Slovák, and Peter W. Michor. Natural Operations in Differential Geometry. Springer, 1993.

  22. [22]

    Saunders Mac Lane and Ieke Moerdijk. Sheaves in Geometry and Logic: A First Introduction to Topos Theory. Universitext. Springer, 1992. Corrected reprint 1994. Internal language: Chapter VI.

  23. [23]

    Jaehoon Lee, Lechao Xiao, Samuel S. Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington. Wide neural networks of any depth evolve as linear models under gradient descent. Advances in Neural Information Processing Systems, volume 32, 2019.

  24. [24]

    Peter LeFanu Lumsdaine. Weak ω-categories from intensional type theory. Logical Methods in Computer Science, 6(3:24), 2010.

  25. [25]

    Jacob Lurie. Higher Topos Theory. Annals of Mathematics Studies 170. Princeton University Press, 2009.

  26. [26]

    Per Martin-Löf. Intuitionistic Type Theory. Bibliopolis, Naples, 1984. Notes by Giovanni Sambin.

  27. [27]

    David Jaz Myers. Categorical systems theory. Book draft, last updated September 3, 2023.

  28. [28]

    Juan C. Perdomo, Tijana Zrnic, Celestine Mendler-Dünner, and Moritz Hardt. Performative prediction. Proceedings of the 37th International Conference on Machine Learning, PMLR 119, pages 7599–7609, 2020.

  29. [29]

    J. P. Pridham. Unifying derived deformation theories. Advances in Mathematics, 224(3):772–826, 2010.

  30. [30]

    D. J. Saunders. The Geometry of Jet Bundles. Cambridge University Press, 1989.

  31. [31]

    David I. Spivak. Functorial data migration. Information and Computation, 217:31–51, 2012.

  32. [32]

    The Univalent Foundations Program. Homotopy Type Theory: Univalent Foundations of Mathematics. Institute for Advanced Study, 2013. See especially Chapters 2, 3, and 9.

  33. [33]

    Benno van den Berg and Richard Garner. Types are weak ω-groupoids. Proceedings of the London Mathematical Society, 102(2):370–394, 2011.

  34. [34]

    Vladimir Voevodsky. A very short note on homotopy λ-calculus. Unpublished note, 2006. Revised version dated 2009 in the author’s archive.

  35. [35]

    Norbert Wiener. Cybernetics: Or Control and Communication in the Animal and the Machine. MIT Press, second edition, 1961. First edition 1948.

  36. [36]

    Greg Yang. Tensor programs II: Neural tangent kernel for any architecture. Preprint, 2020.

联系我们 contact @ memedata.com